Code-only heuristic policies can beat frontier models on Craftax, evals researcher says
JoshPurtell · x · 2026-09-20
In a NetHack benchmark discussion thread, evals researcher Josh Purtell shares a counterintuitive finding: across model sizes on Craftax, he's found that code-only heuristic policies can outcompete models like luna xhigh emitting raw actions—making him hesitant about arbitrary harness modification.
The thread also surfaces an unmeasured paradigm: treating "experience" as a fundamentally different axis from test-time compute. Benchmarks could plot experience (number of observations seen) against performance, the way LLM benchmarks plot test-time compute spend against performance—prompting speculation about BALROGv2.
More from Models
- User searching a game name gets jump-scared by ChatGPT — and it won't stop — ProfessionalRing4307 · 2026-09-20
- Qwen 27B one-shot prompt builds three playable Super Mario clones — EcstaticDentist · 2026-09-20
- Anthropic distillation math: $500M buys ~516T tokens at Opus blended $0.97/M — zephyr_z9 · 2026-09-20
- Jev Classifies 1.6k Bookmarks in 22s, 155x Faster and 10x Cheaper Than GLM 4.7 Flash — iannuttall · 2026-09-20
- GLiNER2 Author Pushes Back on Jev Hype, Highlights Open GLiGuard Guardrails Model — philipvollet · 2026-09-20
- Decision Model Jev Goes Free on Venice API: Typed Answers, No Prose, No JSON Wrangling — 0xAllen_ · 2026-09-20