Celeris-1 Magnus: New Model Claims Top Spot on τ³-bench for Agentic Work
timshi_ai · x · 2026-09-01
Celeris introduced Celeris-1 Magnus, a hybrid diffusion model built specifically for agentic workflows derived from Qwen3.8-27b. On the τ³-bench banking suite (97 tasks), Magnus achieved a 41.2% solve rate with a 55-second median time, outperforming GPT-5.6-sol (38.1%) in both speed and accuracy. It features a "reasoning dial" that boosts performance by 13.4 points when enabled. The API is OpenAI-compatible, allowing seamless integration for existing agents.
More from coding & agent
- Developer Critiques MCP Spec: Overcomplicated JSON-RPC? — PaulMorel · 2026-09-01
- VibeKit MCP Server Manages Deployments, Logs, and Headless Coding — modelcontextprotocol · 2026-09-01
- Distributed.systems Launches Auditable Agent Infrastructure — arthurcolle · 2026-09-01
- How to Stop Context Window Bottlenecks in Data-Heavy MCP Servers — JuicerSocial · 2026-09-01
- Grok Bots Turn LLM Citations Into an SEO Loop for AI Search Ranking — rohanpaul_ai · 2026-09-01
- Search configuration impacts agent accuracy 40x more than model choice — RichardSocher · 2026-09-01