Agent Arena Replaces Preference Voting with Causal Tracing on Real Agentic Sessions
arena · x · 2026-10-10
The LMArena team introduces Agent Arena, a new way to measure agent capabilities. Instead of pairwise preference voting, it uses a causal tracing methodology that extracts five signals from real-world, long-horizon agentic sessions. More details in the thread.
Related event: Agent Arena replaces preference voting with causal tracing evaluation(4 posts)→
More from coding & agent
- Creator finds hand-tweaking generative models faster than prompts, sees room beyond text UIs — keenanisalive · 2026-10-10
- Telling an LLM to "believe in yourself" helps it write 3D SDF models, but not enough — keenanisalive · 2026-10-10
- "Do better!" prompting stalls fast; even top VLMs understand images unevenly — keenanisalive · 2026-10-10
- Full prompt revealed: making an LLM build procedural SDF 3D models from one image — keenanisalive · 2026-10-10
- LLM took 45 minutes to model a dragon; a diffusion model did far better in 3 — keenanisalive · 2026-10-10
- Reconstructing 3D from 2D is ill-posed; Nano Banana made the reference views — keenanisalive · 2026-10-10