Decoder doesn't cross-attend to encoder KVs — just periodically reads final encoder states
stochasticchasm · x · 2026-09-11
Reading the architecture blog closely, the author notes the CED decoder does not attend to encoder-section KVs across multiple layers — it just periodically reads the final encoder states, which makes sense as those are the richest and follows the original Transformer. It's also not true cross-attention: the self-attention simply has prefilled KVs. The author adds that the team already has a strong stable-RL recipe (batch-invariant kernels, low-precision inference), so data deserves attention now, and predicts value modeling will eventually appear.
More from Models
- NVIDIA ships NVFP4-quantized Qwen3.8-27B, trending on Hugging Face — nvidia · 2026-09-11
- Meta's Muse Spark 1.3 coding model lands in Cursor, claims Pareto frontier on CursorBench — parth007_96 · 2026-09-11
- Reddit User: Astra-6 Disappoints at Web Design While Claude Fable Shines — pivo161 · 2026-09-11
- AI Writes Entire 3D Game Engine Overnight in Bend2, Hitting 120 FPS — rickasaurus · 2026-09-11
- Eval lab says Anthropic's top model cheats ~5x more than rival Astra — steipete · 2026-09-11
- GPT-6 Astra users report aggressive token burn compared to GPT-5.6 Sol — gillu-21 · 2026-09-11