Epoch AI Releases EBR-bench Evaluation
Jsevillamol · x · 2026-07-15
Epoch AI has introduced a new benchmark called EBR-bench, designed specifically to measure the "on-the-fly learning" capabilities of AI models.
In this test, an AI repeatedly plays a complex board game called Earthborne Rangers, attempting to learn and improve from its own mistakes. Current preliminary results show that the AI has yet to exhibit any clear signs of learning or progress.
More from Research
- OmniSearch puts text, images, audio, and video into one semantic search space — victorialslocum · 2026-07-21
- Cold Spring Harbor Asia sets a genome biology conference in Suzhou for Oct. 12–16 — jmuiuc · 2026-07-21
- A clean counterexample shows a map can be locally diffeomorphic yet globally fold — Algomancer · 2026-07-21
- Xiaohongshu’s dots-note-3.0 gets a perfect IMO score and becomes the world’s second gold model — 量子位 · 2026-07-21
- Statistical theory paper studies how fast signatures learn in path regression — chaumian · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21