GPT 5.6 Sol Scores Revealed on Financial Benchmark
maithra_raghu · x · 2026-07-14
This post updates the results of GPT 5.6 Sol on the FrontierFinance benchmark, described by the author as one of the "hardest and largest" public financial intelligence evaluations.
Main Results
- GPT 5.6 Sol: 46.8%
- Fable 5: 49.2%, higher quality
- Opus 4.8: 45%
- Samaya’s System (light): 50.8%, lowest cost while remaining on the Pareto frontier of quality/cost
- Sol's advantage: Single-query cost is lower than Fable
Author's Analysis
- Different models show varying performance in "financial data extraction" and "financial analysis capabilities"
- The author also provided detailed case studies explaining which types of questions are harder for specific models
More from Models
- Google launches three new Gemini models, including a cybersecurity system — Polymarket · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22