IndicBankBench: 799-case benchmark shows banking AI assistants top out at 58.2% strict reliability

NPCI · hf · 2026-09-28

NPCI researchers released IndicBankBench, a 799-case benchmark evaluating safety and reliability of language-model assistants in Indian retail banking, spanning 5 operational domains, a capability/refusal domain, and 20 primary axes — with cases, a mock environment, and evaluation harness open-sourced.

Design

Findings

Original post →

More from coding & agent

coding & agent channel →