Benchmark: Deterministic Financial Engine Hits 100%, LLM Pipeline Fails 71%

MuhammadMujtaba21 · reddit · 2026-08-21

The author benchmarked a deterministic verification engine for AI-generated financial claims using 66 test cases. The engine passed 100% (66/66) of cases with pre-defined structured inputs, but the pass rate dropped to 29% (19/66) when using Azure OpenAI GPT-5.1 to generate live claims. Failures were concentrated in claim binding and pipeline execution rather than deterministic calculations or rule application. This suggests the verification architecture works, but the translation layer converting LLM outputs to structured claims needs improvement. The author plans to add granular diagnostics comparing expected claims, raw outputs, and normalized claims.

Related event: Deterministic Financial Validation Engine Hits 100% While LLM Interface Passes Only 29%(2 posts)→

Original post →

More from Research

Research channel →