Primer Tops BigFinanceBench, Reveals 12% Error Rate in Benchmark's Reference Answers

rohanpaul_ai · x · 2026-08-06

Primer published a detailed technical blog breaking down the performance of its finance AI agent on the BigFinanceBench benchmark. Primer achieved 79.1% accuracy, beating top general-purpose models (56%).

Beyond demonstrating the power of their agent harness, the team's manual review uncovered that 12% of the benchmark's public questions contained errors in their reference answers or underlying data. This suggests that current benchmark scores might partly measure how well a system matches flawed references rather than its actual financial analysis capabilities.

Related event: Financial AI Test: Agent Architecture Outshines Base Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →