Pareto 26.9 routes across models, ties GPT-6 on 30 agent tasks at 1/3 the cost
omarsar0 · x · 2026-09-25
- Composio benchmarked 6 models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash.
- Pareto 26.9 from TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer; it tied GPT-6 Astra for first place at roughly 1/3 the cost per successful task, and finished faster than DeepSeek V4 Pro and GLM 5.3 Flash.
- GPT-6 Sol matched Opus 5.5's score while running faster and costing about a quarter as much per successful task.
- Takeaway: blended multi-model routing is emerging as a serious strategy for agent workloads.
More from Models
- Dhravya Shah Details How typesafeAI's Jev Improves AI Memory and Context Engineering — blaizedsouza · 2026-09-25
- Community Speculates DeepSeek V4 Coming Soon After Holiday Post From Liang Wenfeng — teortaxesTex · 2026-09-25
- NaceAI launches Drex, a sub-6B decision model that tops the public Decision Index at 51.73 — ordax · 2026-09-25
- Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind — ivan_bezdomny · 2026-09-25
- Uncensored local model Bonzai 2 27B tops benchmarks, runs on 12GB VRAM — alexcovo_eth · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25