BALROG contributors float a new eval axis: performance vs. experience
mitrma · x · 2026-09-20
Contributors to the BALROG game-agent benchmark discuss a new evaluation dimension: just as LLM benchmarks plot test-time compute spend against performance, this paradigm could plot experience (number of observations seen) against performance. If you have the sim, you can query it and roll that into the test-time-compute measure—an axis they call currently unmeasured, hinting at BALROGv2.
The thread also proposes a bigger idea: to make progress on NetHack and games in general, models should be allowed to tune their harnesses quite freely—incorporating wiki knowledge and building custom tools that amortize low-level reasoning and control.
More from Research
- New research traces distillation length inflation to student-teacher EOS token mismatch — tw_killian · 2026-09-20
- XGEN Labs unveils generative world simulation JING+DAO, tops WBench leaderboard — hey_abusiddik · 2026-09-20
- Schmidhuber: LLMs aren't truly creative because they lack compression progress — SchmidhuberAI · 2026-09-20
- Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline — This_Cell_1829 · 2026-09-20
- François Fleuret nicknames his training curves; researchers admit they curse baselines too — giffmana · 2026-09-20
- HN: How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — petrusenko_max · 2026-09-20