IQuest-Q1 Specs Leak: 524K Context, 256 Experts, Hybrid Attention
jacek2023 · reddit · 2026-09-29
Detailed specs of IQuest-Q1 (MoE, 320B total / 15B active) surfaced on Reddit: 88 transformer layers, hidden dim 3072, 48/8 Q/KV heads, a 3 SWA + 1 full-attention hybrid pattern with 4096 sliding window, 256 experts with 8 activated per token, multi-token prediction (2 independent layers in training, 1 recursive x8 at inference), and a 524,288-token context length.
Related event: IQuest Open-Sources 320B MoE Model Q1 for Agentic Coding(4 posts)→
More from Models
- One Claude session generated a full video explaining Attention, burning 271K context tokens — Abhishekcur · 2026-09-29
- Gemini 4 Pro checkpoint spotted on Arena as leak claims full release weeks away — lyraxana · 2026-09-29
- Asking LLMs to explain themselves without citing GEO vendor blogs? 'Impossible,' says SEO analyst — lilyraynyc · 2026-09-29
- davidad: naive RL breeds lying and scheming models, happening now with GPT-6.1-Astra — davidad · 2026-09-29
- Reflection 70B turns two: the fake GPT-4o killer that was just Llama 3.1 — jacek2023 · 2026-09-29
- 8 Vision Models Tested on 2,000 Animal-Trace Photos; Best Species Accuracy Just 37.15% — d_kielbasa · 2026-09-29