FrontierFinance Financial Agent Benchmark Released
maithra_raghu · x · 2026-07-09
Samaya has released FrontierFinance, an open-ended AI agent benchmark covering the entire investment lifecycle to evaluate scenarios like screening and discovery, company research, industry/macro analysis, earnings events, and coverage monitoring. The benchmark includes 220 samples and 11,543 expert rubrics, providing comparative results for models like Claude, GPT, Gemini, GLM, and DeepSeek.
Related event: Samaya Releases FrontierFinance Benchmark for Financial AI Agents(7 posts)→
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22