Gate the action, not the confidence score: why agent risk control should follow capabilities
arslannasir128 · reddit · 2026-09-26
The author argues confidence floors are the least trustworthy control in agent setups: teams tune the number for weeks, yet costly failures still slip through—and a confident failure never raises its hand, hiding in logs nobody reads while averages look healthy.
His practical rule: draw the line by what a call does, not how the model feels. Reads stay safe even when the model is wrong; anything moving money or canceling accounts waits for a human click. He closes by asking whether others gate by capability, dollar amount, or a score they trust.
More from coding & agent
- Using Opus 5.5 to make a video explaining Algorithm W and Hindley-Milner type inference — ctjlewis · 2026-09-26
- Runway MCP brings Gen-4.5 video and image generation into Claude chats — runwayml · 2026-09-26
- One 'Harmless' Agent Permission Turned Into a Backdoor for Everyone — Adorable-Algae6903 · 2026-09-26
- DEV·TV: a single-HTML-file TV playing Hugging Face models and AI papers hits Product Hunt — TradingCardGirl · 2026-09-26
- MCP Wallet Agents Need Boring Permission Rules Before Going Agentic — Any-Sir-7622 · 2026-09-26
- Anthropic engineer: Claude Tag writes >50% of my PRs every day — bcherny · 2026-09-26