ByteDance EdgeBench: Evaluating AI Agents' 10-Hour Self-Iteration Capabilities
karminski3 · x · 2026-07-06
ByteDance's Seed team released the EdgeBench evaluation framework, designed specifically to assess AI agents' ability to autonomously iterate and improve over long horizons (up to dozens of hours) in a real, runnable environment. Unlike single-task evaluations, it focuses on whether an agent can leverage environmental feedback, adjust strategies based on failures, and continuously enhance performance.
The tester used a locally run 35B-A3B model (4-bit quantization) to tackle the classic Roguelike game DCSS (command-line version). The task required the Agent to write Lua bot scripts via CodeX CLI to automatically explore, fight, and descend levels within roughly 70 minutes. The initial automated score was 15.2, which improved to a maximum of 33.4 after multiple iterations, clearly showcasing the AI's true learning trajectory in long-horizon tasks.
Related event: ByteDance Releases EdgeBench to Evaluate Long-Horizon Agent Evolution(4 posts)→
More from coding & agent
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11