Should models tune their own harnesses? NetHack eval discussion says maybe
mitrma · x · 2026-09-20
Discussing why models struggle at NetHack, an evals contributor argues that to make progress (in NetHack and in general) we should let models tune their harnesses quite freely—incorporating wiki knowledge and building custom tools that amortize low-level reasoning and control.
A BALROG contributor notes that BALROG measures zero-shot performance, quipping: who has ever beaten NetHack on their first attempt, even with wiki usage?
More from Research
- New research traces distillation length inflation to student-teacher EOS token mismatch — tw_killian · 2026-09-20
- XGEN Labs unveils generative world simulation JING+DAO, tops WBench leaderboard — hey_abusiddik · 2026-09-20
- Schmidhuber: LLMs aren't truly creative because they lack compression progress — SchmidhuberAI · 2026-09-20
- Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline — This_Cell_1829 · 2026-09-20
- François Fleuret nicknames his training curves; researchers admit they curse baselines too — giffmana · 2026-09-20
- HN: How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — petrusenko_max · 2026-09-20