Open 33B multimodal Agnes-3.0-Flash ships hybrid delta-rule attention with 262K context
Skyline34rGt · reddit · 2026-09-12
Agnes-AI released Agnes-3.0-Flash, an open 33B multimodal model (Artificial Analysis score 36, which ironically lists it as 'proprietary'). It uses a hybrid-attention decoder: 3 of every 4 layers are gated delta-rule recurrent (state independent of sequence length), only 18 of 72 layers hold a growing KV cache.
Key specs: 262,144-token context, 248,320 vocab, 6:1 GQA global attention, fp32 recurrent state, SwiGLU FFN with parallel branch, 3-axis mrope, a 27-layer vision tower for text/image/video understanding, adjustable reasoning effort and tool calling.
More from Models
- Reddit user argues 20B-32B dense / A2B-A4B MoE is the sweet spot for perfect models — Robert__Sinclair · 2026-09-12
- User reports ChatGPT memory has degraded: forgets facts it claims to save — PerilousParanoia · 2026-09-12
- Devs split on GPT-6 Astra: fast and token-cheap but code feels 'alien'; costly stack proposed — beffjezos · 2026-09-12
- After weeks of testing: Sonnet for clean docs, Opus when structure must be inferred — SpryShade · 2026-09-12
- DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s — AccBalanced · 2026-09-12
- Codex reset rolling out now, and GPT-Image 2.5 ships a new sketch feature — koltregaskes · 2026-09-12