Gemini 3.6 Flash Tops FrontierFinance Benchmark at $2.41/Query
tulseedoshi · x · 2026-07-31
In the FrontierFinance open benchmark for finance AI agents by Samaya AI, Gemini 3.6 Flash scored 46.3%, beating Claude Opus 4.8 and tying with GPT-5.6 Sol.
While delivering frontier-level accuracy, the model is highly efficient: it costs only $2.41 per query and runs in 164 seconds, making it cheaper and faster than peers of the same tier.
Analysis indicates that Gemini 3.6 Flash is notably better than Opus 4.8 at surfacing qualitative and contextual insights. It also leads in Screening & Discovery, the benchmark's hardest use case, making it ideal for financial applications requiring a balance of speed, cost-efficiency, and high performance.
More from Models
- Users Report Claude Opus Frequently Lies Badly in Interactions — rickasaurus · 2026-07-31
- OpenAI Accused of Cherry-Picking Data in Benchmark Graphs — ns123abc · 2026-07-31
- Agent Arena Leaderboard Updates: Claude Fable 5 Takes #1 — arena · 2026-07-31
- Sol-5.6 Ultra Mode Reported to Overthink and Get Stuck in Loops — AIandDesign · 2026-07-31
- Claude API Retains Thinking Blocks by Default; Developers Urge OpenAI to Follow Suit — steipete · 2026-07-31
- Debate Erupts Over OpenAI vs Anthropic Default Chain-of-Thought Retention in APIs — steipete · 2026-07-31