LMArena Benchmarks: Open-Weight Hy3 Ranks #2 in Frontend Code Arena
arena · x · 2026-07-22
LMArena released the latest benchmark results for the Frontend Code and Agent Arenas. The open-weight model Hy3 performed impressively, ranking #2 among open-weight models (#16 overall) in the Frontend Code Arena, and made the top 20 in reference-based design, gaming, and content creation tools.
However, in the Agent Arena based on over 8,000 live agentic sessions, Hy3 ranked #25 overall with a -2.2% net improvement. It showed declines in confirmed success and steerability, though slight improvements were noted in Bash recovery and tool hallucination control.
More from coding & agent
- Vercel AI SDK v7 Support Ships in Agents Framework — threepointone · 2026-07-24
- Testing Infinity: Agent Autonomously Spawns Claude to Write and Steer Code — jasonkneen · 2026-07-24
- APL Point-Free Style May Revive in the Era of AI Coding Agents — satnam6502 · 2026-07-24
- Point-free code may look better once machines write and verify it — satnam6502 · 2026-07-24
- Use a second model to approve tool calls and review an agent’s trajectory — corbtt · 2026-07-24
- PRO-LONG gives LLM agents searchable programmatic memory and lifts ARC-AGI-3 by 18 points — dair_ai · 2026-07-24