Same weights, dumber model: 100k-token logit study exposes inference stack pitfalls

量子位 · wechat · 2026-08-29

Level1Techs user thr3e ran rigorous experiments capturing full logits across 100k+ tokens with Qwen3.6-27B on an RTX PRO 6000 Blackwell, explaining why local LLMs often feel dumber than official ones.

The vLLM nightly container carries 734 packages; the author warns that low KL-divergence numbers on model cards are uninterpretable without full disclosure. Test tooling is being packaged for distribution.

Original post →

More from Infra

Infra channel →