FinFIRST details: 123 expert tasks, 701 atomic criteria and 12,300 rubric points for financial search agents
FellMentKE · x · 2026-09-04
Author adds details on FinFIRST, the financial search agent benchmark Ant Group built with professional support from CICC's investment banking team. Version 1 includes 123 expert-authored tasks, 701 atomic criteria and 12,300 rubric points. The launch claim and evaluation framework answer different questions: one gives people a model to adapt, the other gives researchers a detailed way to examine where agents succeed, fail, or need closer review.
Related event: Ant Group Open-Sources FinFIRST Benchmark for Financial Search Agents(2 posts)→
More from Models
- OpenAI ships GPT-6 Astra; Brockman suggests it could be AGI — steph_palazzolo · 2026-09-04
- ChatGPT desktop update hides Ultra by default, adds model-then-reasoning picker — koltregaskes · 2026-09-04
- Meta's Muse Spark 1.3 ranks second-best agentic model at ~1/3 cost — ryanshrout · 2026-09-04
- Zvi opens reaction thread on Mythos+Fable 5.1, 'world's most advanced AI' — TheZvi · 2026-09-04
- User burns through Kimi free quota via daily automations, weighs which subscription to buy — Inevitable_Fold_9081 · 2026-09-04
- Meta's Muse Spark 1.3: 38% fewer errors than Opus 5 at $0.85 per correct task — ryanshrout · 2026-09-04