Third-party AI evaluation faces gaps in expertise and consequences

Practitioners highlight flaws in third-party AI evaluation: bank risk-model assessors could not even pose meaningful questions about gradient-boosted trees, and an analyst asks what actually happens when evaluators of Anthropic's systems find unacceptable issues.

2026-09-13 ~ 2026-09-13 · 2 related posts