Dev audits own SOC agent, finds its 96% confidence score was fake — three bugs
Life_Rest_2488 · reddit · 2026-09-30
A developer audited their own incident-response agent (FastAPI + Next.js, Hindsight memory, Groq) and found the dashboard showed 96% confidence for everything, caused by three bugs in the scoring layer:
- Similarity was a rank: similarity = round(0.95 - idx 0.04, 2) — percentages were just list positions, and every match was hard-coded as a "success".
- "Times Deployed" was recall size: timesused = max(len(results), timesused) made every success rate (N−2)/N.
- Confidence saturated: a 0.4/0.6 blend of fake similarity and fake success rate always lands near 0.96, and a max(…, 0.50) floor meant the agent could never admit it knew nothing.
Fixes: counters start at zero and update only from analyst feedback; similarity comes from the recall payload or isn't shown; remove the floor so cold start looks like cold start. Playbook selection remains a keyword lookup. Takeaway: any percentage in a UI should be traceable to the line of code that produced it.
More from coding & agent
- Dev adds MCP interface to exe, hooks it into ChatGPT to do real work from his phone — davidcrawshaw · 2026-09-30
- Theo's own Terminal Bench 4 runs: GPT-6.1 Sol beats Opus 5.5 at ~1/30th the price — dkundel · 2026-09-30
- Agent team pattern improves inter-agent communication and deep exploration — omarsar0 · 2026-09-30
- GPT-6.1 Sol ran in circles for days on ts-rust, sparking regression complaints — karmay007 · 2026-09-30
- Developer Gets Grok Bot Running Inside a Robotaxi and Opening a PR — Baconbrix · 2026-09-30
- OpenAI launches Decisions API powered by GPT-6 Luna for real-time app decision-making — OpenAIDevs · 2026-09-30