Code-only heuristic policies can beat frontier models on Craftax, evals researcher says

JoshPurtell · x · 2026-09-20

In a NetHack benchmark discussion thread, evals researcher Josh Purtell shares a counterintuitive finding: across model sizes on Craftax, he's found that code-only heuristic policies can outcompete models like luna xhigh emitting raw actions—making him hesitant about arbitrary harness modification.

The thread also surfaces an unmeasured paradigm: treating "experience" as a fundamentally different axis from test-time compute. Benchmarks could plot experience (number of observations seen) against performance, the way LLM benchmarks plot test-time compute spend against performance—prompting speculation about BALROGv2.

Related event: BALROG NetHack benchmark debate: harness self-modification and degenerate solutions(7 posts)→

Original post →

More from Models

Models channel →