BALROG NetHack results questioned: worse than last week's numbers?
mitrma · x · 2026-09-19
In the BALROG NetHack eval discussion, a participant questions creusroger's posted results: were they run with a default harness, and why do they seem worse than the results he was sharing last week? The point is made that BALROG assesses zero-shot performance—but who has ever beaten NetHack on their first attempt, even with wiki usage?
More from Research
- New research traces distillation length inflation to student-teacher EOS token mismatch — tw_killian · 2026-09-20
- XGEN Labs unveils generative world simulation JING+DAO, tops WBench leaderboard — hey_abusiddik · 2026-09-20
- Schmidhuber: LLMs aren't truly creative because they lack compression progress — SchmidhuberAI · 2026-09-20
- Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline — This_Cell_1829 · 2026-09-20
- François Fleuret nicknames his training curves; researchers admit they curse baselines too — giffmana · 2026-09-20
- HN: How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — petrusenko_max · 2026-09-20