OrcaRouter stress-tests JEV: dropping autoregressive decoding could cut inference cost 10-100x
Dan_Jeffries1 · x · 2026-09-23
OrcaRouter spent days reproducing, breaking, and improving JEV internally. Key takeaways:
- The core insight is probably right: if the answer space is bounded, don't generate token by token — removing autoregressive decoding can cut inference work by 1–2 orders of magnitude.
- Data, not RLCD, is the moat: Laya already open-sourced the implementation and weights; the missing piece is the synthetic data recipe, and OrcaRouter's experiments point the same way.
- "Open source already beat JEV" is a benchmark illusion: the same checkpoint scores 0.769 in-distribution but only 0.541 OOD — the apparent breakthrough largely disappears under a different distribution.
- Compute-optimal ≠ learnability-optimal: moving state outside the problem sequence to save FLOPs cost 27 points.
More from Models
- GPT-6 Sol accused of ignoring configured skills that work fine on GPT-5.6 — Aber-so-richtig · 2026-09-23
- Opus 5.5 ultra wows with 4-agent, 90-minute fully-from-scratch build — repligate · 2026-09-23
- No Kimi release this week, says leaker: Moonshot staff go on holiday Friday — ChrisGPT · 2026-09-23
- Screenshot suggests GPT-5.6 Sol is free for ChatGPT Go subscribers on Windows — sam619007 · 2026-09-23
- Opus 5.5 wins the day as GPT-6 Sol skips ChatGPT, sparking model debate — koltregaskes · 2026-09-23
- Opus 5.5 becomes a daily driver: faster, cheaper than Opus 5, plus reset credits — addyosmani · 2026-09-23