omarsar0 on custom agent benchmarks: public numbers don't mean much, Jev Router rivals GPT-6 Astra low
omarsar0 · x · 2026-09-27
AI researcher omarsar0 shares a self-built agent harness experiment, concluding his numbers are promising but may not generalize to every task.
- Key lesson: public benchmark numbers don't mean much — what works for you may not work for others
- Theo's scoped DeepSWE tests show Jev Router performance comparable to GPT-6 Astra low, at the cost of longer runs
- The two tests measure different things, with both disagreements and agreements
- He also finds reasoning effort an interesting variable to experiment with
Related event: Researchers Call for Better Benchmarks as Model Routing Faces Criticism(2 posts)→
More from coding & agent
- Opus 5.5 for motion graphics: dev builds a full template app, shares Midjourney prompts — techhalla · 2026-09-27
- Devs Are Just Yes Machines Now: Complaining About Claude Extension's Constant Prompts in VS Code — ThePeterMick · 2026-09-27
- Routing Evals Are Missing, Says Elvis Saravia as OpenRouter Criticism Rages Without Evidence — omarsar0 · 2026-09-27
- Opus 5.5 prompt turns code-rendered frames into an epic 'History of Chinese Civilization' film — dotey · 2026-09-27
- Meta's Muse gives every user their own VM, called the most careful consumer agent sandbox yet — AccBalanced · 2026-09-27
- SafeScript: a Turing-incomplete JS subset lets agent policies replace code review — uriwa · 2026-09-27