FinanceComplexQA adds a 2,026-task benchmark for agentic reasoning on financial docs
Beihang · hf · 2026-07-24
FinanceComplexQA benchmarks agentic reasoning on financial documents
This paper introduces FinanceComplexQA, a benchmark for open-ended reasoning over industrial financial documents.
- The authors first design Finance-LaTeX SKILL, a skill for synthesizing complex financial documents from expert knowledge.
- Using an agent workflow built on that skill, they generate 2,000 professional financial documents and 6,000 QA pairs.
- The benchmark contains 2,026 deep-research tasks covering 1,009 financial documents.
- It supports bilingual evaluation and spans six mainstream scenarios and seven tasks.
- Evaluation is done with an Agent-as-a-Judge setup and multiple metrics.
- The paper uses the benchmark to test leading RAG systems and agentic reasoning tools on numerical computation, multi-hop reasoning, summarization, and industry analysis.
- The authors also analyze failure cases to show where current systems still struggle on real financial documents.
The main contribution is a more realistic benchmark for document-heavy financial QA, plus a dataset and evaluation setup that stress agentic reasoning rather than simple retrieval.
More from coding & agent
- Builder rebuilds a Claude Code software factory around planning, implementation, review — blaizedsouza · 2026-07-24
- Default Codex CLI with GPT-5.5 scores 92.3% on XBOW, sparking benchmark fatigue — moyix · 2026-07-24
- Sebastian Raschka will discuss DeepSeek-V4, GLM-5.2 and open-weight coding agents — hugobowne · 2026-07-24
- Antigravity CLI 1.1.6 makes custom agents editable as Markdown files — rseroter · 2026-07-24
- Microsoft Research’s ReOPD reuses teacher prefixes to distill multi-turn agents offline — MicrosoftResearch · 2026-07-24
- Compound Engineering 3.20 splits AI coding across multiple models and adds handoff snapshots — danshipper · 2026-07-24