Claude's fraud detector scored F1 0.87 — until data leakage was fixed and it fell to 0.70
hugobowne · x · 2026-09-29
hugobowne, twiecki and lfiaschi86 asked Claude to build a fraud detector and uncovered a classic eval trap:
- Great-looking numbers: F1 of 0.87 and ROC AUC of 0.99 — but the evaluation was invalid.
- Two flaws: Claude split transactions randomly across time (leaking future data) and used a planted feature acting as a proxy for the fraud label, so nothing predicted real-world performance on new transactions.
- After fixing: with a temporal holdout and the leaking feature removed, F1 dropped to 0.70; recall on the highly connected transactions that mattered most was just 0.21.
- Verification playbook: check whether the data split matches deployment, whether features exist at prediction time, and whether performance holds on the transactions that matter. Agents can run checks, but humans judge the evidence.
More from coding & agent
- Fireworks Launches FireRouter: 98.1% of Opus Accuracy at 57% Lower Coding Cost — nicolechirps · 2026-09-29
- "How often can a hallucination become a real action?" Reddit debates production agent safety — Informal-Dust4499 · 2026-09-29
- liminal_groupchat: open-source AI group chat with cross-thread persistent memory — liminal_bardo · 2026-09-29
- Claude Code Projects defaults to low effort, and an Anthropic engineer teases Sonnet 5.5 — lydiahallie · 2026-09-29
- AgentCribs SF set for Oct 6 with OpenClaw creator steipete in fireside chat — steipete · 2026-09-29
- Munder Difflin open-sources a multi-agent harness turning Claude Code into a 24/7 agent office — chaitanyagiri · 2026-09-29