Better Representations Yield Large Gains on NetHack and Craftax RL Benchmarks
GlenBerseth · x · 2026-09-28
Roger Creus shared recent work showing that with the right learned representations, reinforcement learning can achieve much better results on NetHack and Craftax, two notoriously hard exploration benchmarks. Glen Berseth highlighted the work and suggested updating previously reported numbers.
More from Research
- Estimate: crudely describing human biology needs 1000x more data than humanity stores — IgorCarron · 2026-09-28
- SkillGym fine-tuning lifts Qwen3.5 35B past Claude Sonnet 4.6 on agentic coding benchmarks — dair_ai · 2026-09-28
- RL's limit: without a scoring function there's no signal, and 'learning' is inflated jargon — gerardsans · 2026-09-28
- Information geometry exactly characterizes Chernoff info between Gaussians without closed forms — FrnkNlsn · 2026-09-28
- d9bench scores decision models on how fair their 9-sided dice rolls are — swishfever · 2026-09-28
- LLMs can reconstruct documents from structural metadata alone, engineer finds — ChuckDBrooks · 2026-09-28