IQuest-Q1 Specs Leak: 524K Context, 256 Experts, Hybrid Attention

jacek2023 · reddit · 2026-09-29

Detailed specs of IQuest-Q1 (MoE, 320B total / 15B active) surfaced on Reddit: 88 transformer layers, hidden dim 3072, 48/8 Q/KV heads, a 3 SWA + 1 full-attention hybrid pattern with 4096 sliding window, 256 experts with 8 activated per token, multi-token prediction (2 independent layers in training, 1 recursive x8 at inference), and a 524,288-token context length.

Related event: IQuest Open-Sources 320B MoE Model Q1 for Agentic Coding(4 posts)→

Original post →

More from Models

Models channel →