Andriy Burkov on why Jev-style LLM calibration can't beat an LLM at real probabilities
burkov · x · 2026-09-23
ML author Andriy Burkov posted a technical critique of Jev:
- Calibration isn't ground truth: You can RL-finetune an LLM using the gap between softmax probabilities and "real" probabilities from training data (what people call calibration), but the model still won't learn to output true probabilities for arbitrary questions — ground-truth probabilities can't exist for every possible question, from "civilian or military?" to "buy or sell this stock?" to which sock to put on first.
- Speed implies a quality ceiling: An LLM's time-to-first-token reflects the compute spent on a decision. If Jev emits its full output faster than an equal-sized LLM prints its first token, it has automatically lost on quality.
His conclusion: such fast, small decision models gain speed precisely by giving up probabilistic reliability on open-ended questions.
Related event: Burkov Dismisses Mysterious Project Jev's Speed Claims as BS(5 posts)→
More from Models
- GPT-6 Sol and Luna already usable in Codex, early user reports — airesearch12 · 2026-09-23
- GPT-6 Sol scores slightly below GPT-5.6 Sol on DeepSWE, only cheaper — Angaisb_ · 2026-09-23
- Matt Shumer on Opus 5.5: 'feels like a much smarter Opus 4.6' — mattshumer_ · 2026-09-23
- GPT-6 Sol claimed to cost 50% less than GPT-5.6 Sol — cedric_chee · 2026-09-23
- Four frontier models in days: Grok 4.7, Opus 5.5, GPT-6 Sol and Luna — msg · 2026-09-23
- OpenAI reportedly rolling out GPT-6 Sol and GPT-6 Luna on ChatGPT, Codex, and APIs — testingcatalog · 2026-09-23