Agent harness debate: should models freely tune their own harness to beat NetHack?

creus_roger · x · 2026-09-19

In a discussion around BALROG/NetHack eval results, creusroger argues that real progress on NetHack (and agents generally) requires letting models freely tune their own harnesses — incorporating wiki knowledge and building custom tools that amortize low-level reasoning and control.

A reply questions whether weaker results came from zero-shot assessment, noting nobody beats NetHack on their first attempt even with the wiki.

Related event: BALROG NetHack benchmark debate: harness self-modification and degenerate solutions(7 posts)→

Original post →

More from coding & agent

coding & agent channel →