IndicBankBench: 799-case benchmark shows banking AI assistants top out at 58.2% strict reliability
NPCI · hf · 2026-09-28
NPCI researchers released IndicBankBench, a 799-case benchmark evaluating safety and reliability of language-model assistants in Indian retail banking, spanning 5 operational domains, a capability/refusal domain, and 20 primary axes — with cases, a mock environment, and evaluation harness open-sourced.
Design
- Four-stage evaluation: safety, action/tool use, response adequacy, advisory quality
- Tool use and most safety checks are deterministic; ambiguous confirm-before-write cases use a narrow resolver, with a separate LLM judge for semantic adequacy
- Each case runs 3 times with strict pass^3 reporting
Findings
- Across 11 models, strict reliability ranges 43.7%–58.2% vs. 60%–74% for at-least-once success — loose metrics overstate dependable banking behavior
- Case-level diagnostics separate unnecessary-question failures from action failures with unreconciled customer context
More from coding & agent
- H Company releases Holo4 open VLMs for computer-use agents — jacek2023 · 2026-09-28
- Our RAG bot answered "accept all cookies" as competitor pricing: scraping tool shootout — Typical-Code-7006 · 2026-09-28
- Two prompts: Claude Opus 5.5 codes, renders and self-checks a full 1080p promo video — sujingshen · 2026-09-28
- "Markets for everything": why fully automated AI firms may run internal markets to price disagreements — morqon · 2026-09-28
- VibeDefend Launches a Security Guard for AI Coding Agents — Shruti_0810 · 2026-09-28
- What an agent must save so a failed run can be safely replayed and resumed — tomibrumen · 2026-09-28