Frontier Models Struggle in Enterprise Insurance Tasks with 20% Drop in Pass Rate

ajratner · x · 2026-08-07

SnorkelAI introduced UNDERWRITE at CAIS, an expert-built, multi-turn benchmark grounded in real enterprise conditions like insurance underwriting to evaluate AI agents.

Across 13 frontier models, results revealed significant gaps between research performance and enterprise readiness. Models exhibited domain hallucination despite having tool access, and pass^k results dropped by 20%.

Original post →

More from coding & agent

coding & agent channel →