Tuned 26B diffusion model beats Jev after adding 512-token decode-time reasoning budget
rickasaurus · x · 2026-09-21
Developer pomterree tested test-time scaling on DJev by adding a 512-token budget to the default-off think knob, letting the model reason before returning a structured read. A tuned 26B diffusion model then trivially beats Jev, which already topped the author's earlier benchmark of major Jev variants.
More from Models
- User reports Astra is faster and more token-efficient, asks OpenAI what changed — McDonaghMatthew · 2026-09-21
- StepFun launches Step 5 Preview: 600B MoE, 27B active, 1M context at 65% lower cost — realmrfakename · 2026-09-21
- Ling 3.0 Tiny vs Gemma 26B-A4B: 3x Smaller VRAM, But Accuracy Halved — autonoma_2042 · 2026-09-21
- Two-Person Lab Open-Sources 27B Writing Model Hemmingway-1, 1330 on EQ-Bench 4 — lukinator644 · 2026-09-21
- Hamel Husain: Calling LLMs Classifiers Is Fine, but 'Just a Classifier' Sells Them Short — HamelHusain · 2026-09-21
- Zero-shot demos are a poor measure of model intelligence, dev argues — brandon_galang · 2026-09-21