Code-only policies beat raw-action agents on Craftax, sparking agent eval design debate
JoshPurtell · x · 2026-09-20
Continuing the NetHack harness debate, JoshPurtell shares eval experience on Craftac across model sizes: code-only heuristic policies can outcompete agents like luna xhigh emitting raw actions.
He warns that if the optimal solution is an agent writing an end-to-end code policy and running it against the env — or interrupting code exec only for a few scenarios — free harness modification may lead to degenerate solutions, questioning the eval design itself.
More from coding & agent
- Accorduon turns a foldable iPhone into an accordion — the hinge is the bellows — rounak · 2026-09-20
- Agents no longer need setup: hand them a bare machine and they fetch everything themselves — vivekhaldar · 2026-09-20
- DiffusionGemma 26B-A4B turned into a local System One fast-decision model via vLLM patch — solyarisoftware · 2026-09-20
- Devin's New SWE 2 Model Goes Free and Unlimited, Claims Kimi K3 Post-Training — silasalberti · 2026-09-20
- Dev uses Claude to label 200 unsupervised visual clusters in a painting-retrieval app — jamievurnilla · 2026-09-20
- Open-source semgrep finds code and logs by meaning, not regex, with cross-language matching — udmrzn · 2026-09-20