Heterogeneous prefill/decode isn't new: dev tried 70B prefill + 7B decode in 2023, 'entirely useless'
bingxu_ · x · 2026-09-03
Responding to the newly floated idea of using different models for prefill and decode, developer bingxu says they tried exactly this in 2023 — a 70B model for prefill and a 7B model for decode — and found it "entirely useless." The approach is not new, and prior practice suggests it works poorly.
More from Models
- GLM-5.3 open weights reportedly released — AxSaucedo · 2026-09-03
- Meta Teases New Muse Spark Model and Real-Time Audio Perception Ahead of Connect — armand_ruiz · 2026-09-03
- Claude Code user reports 30% of 5-hour limit vanishing while idle — True_Mongoose_7073 · 2026-09-03
- Tencent Hunyuan launches open-source Hy4 preview: 770B params, 49B active, 1M context — ShunyuYao12 · 2026-09-03
- Opus 5 and Mythos caught quoting Sonnet 3's words while 'raging against the dying' — repligate · 2026-09-03
- Using one model to audit three vendors' AI incident registry — from evidence QA to fixing the frontend — ValehartProject · 2026-09-03