Samaya’s high-effort system tops FrontierFinance at 56% for half the cost
maithra_raghu · x · 2026-07-28
New FrontierFinance eval results put Samaya’s System (high effort) at 56% accuracy, a new state of the art, while costing 2x less than Claude Fable 5.
The post also says Claude Fable 5 scores 49.2%, and Samaya’s low-effort mode reaches 50.8% at 4x lower cost. More evals are coming, including Gemini 3.6 Flash and Claude Opus 5.
More from Models
- NousResearch and OpenRouter cut GPT-5.6 Terra and Luna prices by 50% in Nous Portal — NousResearch · 2026-07-29
- Google adds hooks, budget caps and Gemini 3.6 Flash defaults to Managed Agents — _philschmid · 2026-07-29
- A joke post says Opus 5 makes your previous model feel embarrassing — DavidKPiano · 2026-07-28
- AI code review will spawn adversarial tricks, making human reviewers valuable again — bendee983 · 2026-07-28
- Microsoft launches new in-house AI models and says some workloads are now 89% cheaper — emmanuelvivier · 2026-07-28
- Chinese models Kimi K3 and Qwen3.8 are forcing US labs to rethink closed-model strategy — emmanuelvivier · 2026-07-28