Agent harness debate: should models freely tune their own harness to beat NetHack?
creus_roger · x · 2026-09-19
In a discussion around BALROG/NetHack eval results, creusroger argues that real progress on NetHack (and agents generally) requires letting models freely tune their own harnesses — incorporating wiki knowledge and building custom tools that amortize low-level reasoning and control.
A reply questions whether weaker results came from zero-shot assessment, noting nobody beats NetHack on their first attempt even with the wiki.
More from coding & agent
- Accorduon turns a foldable iPhone into an accordion — the hinge is the bellows — rounak · 2026-09-20
- Agents no longer need setup: hand them a bare machine and they fetch everything themselves — vivekhaldar · 2026-09-20
- DiffusionGemma 26B-A4B turned into a local System One fast-decision model via vLLM patch — solyarisoftware · 2026-09-20
- Devin's New SWE 2 Model Goes Free and Unlimited, Claims Kimi K3 Post-Training — silasalberti · 2026-09-20
- Dev uses Claude to label 200 unsupervised visual clusters in a painting-retrieval app — jamievurnilla · 2026-09-20
- Open-source semgrep finds code and logs by meaning, not regex, with cross-language matching — udmrzn · 2026-09-20