Shipping a GBT model into a scorecard world: third-party evaluators couldn't even ask questions
neal_lathia · x · 2026-09-13
The author recalls shipping a gradient boosted tree model in a domain that traditionally relies on manually created scorecards — the third-party evaluator had no idea what it was and couldn't ask any useful questions. The real issue, he adds, is what happens when evaluators do find something unacceptable: look at the size of fines some banks have received — for some that would be the end, for others they pay and carry on.
Related event: Third-party AI evaluation faces gaps in expertise and consequences(2 posts)→
More from Companies & People
- Developer grills OpenAI: has AGI been declared, and where is the expert panel? — tomchapin · 2026-09-13
- Anthropic says engineers now ship 8x more code as AI nears recursive self-improvement — bibryam · 2026-09-13
- OpenAI hosts Astra Commons event where top users teach Codex harness and GPT-6 Astra workflows — paw_lean · 2026-09-13
- AI scientist's six-point responsibility manifesto: against weaponization and pre-IPO hype — NandoDF · 2026-09-13
- Founder essay: building something extraordinary requires taking your bubble far too seriously — DominiqueCAPaul · 2026-09-13
- Swiss AI Safety Days 2026 doubles to two days, Stuart Russell to keynote at ETH Zurich — maksym_andr · 2026-09-13