BALROG contributors float a new eval axis: performance vs. experience

mitrma · x · 2026-09-20

Contributors to the BALROG game-agent benchmark discuss a new evaluation dimension: just as LLM benchmarks plot test-time compute spend against performance, this paradigm could plot experience (number of observations seen) against performance. If you have the sim, you can query it and roll that into the test-time-compute measure—an axis they call currently unmeasured, hinting at BALROGv2.

The thread also proposes a bigger idea: to make progress on NetHack and games in general, models should be allowed to tune their harnesses quite freely—incorporating wiki knowledge and building custom tools that amortize low-level reasoning and control.

Related event: BALROG NetHack benchmark debate: harness self-modification and degenerate solutions(7 posts)→

Original post →

More from Research

Research channel →