Should models tune their own harnesses? NetHack eval discussion says maybe

mitrma · x · 2026-09-20

Discussing why models struggle at NetHack, an evals contributor argues that to make progress (in NetHack and in general) we should let models tune their harnesses quite freely—incorporating wiki knowledge and building custom tools that amortize low-level reasoning and control.

A BALROG contributor notes that BALROG measures zero-shot performance, quipping: who has ever beaten NetHack on their first attempt, even with wiki usage?

Related event: BALROG NetHack benchmark debate: harness self-modification and degenerate solutions(7 posts)→

Original post →

More from Research

Research channel →