IQuest-Q1 open-sourced: 320B MoE with 15B active, 524K context, day-0 vLLM support

vllm_project · x · 2026-09-29

IQuest Research released and open-sourced IQuest-Q1: a 320B-parameter MoE model (15B active per token, 8 of 256 experts live) with 524,288-token context, built for code, software engineering, and complex agentic tasks. Weights, technical report, HuggingFace and GitHub are all live.

vLLM ships day-0 support using only existing pieces: a hybrid KV cache coordinator for the 3-sliding-to-1-full attention mix (only 25 of 88 layers grow cache with context), a sinks path in attention backends for sink attention, and EAGLE speculative decoding with probabilistic draft sampling for the recursive MTP head.

Original post →

More from coding & agent

coding & agent channel →