Building a reliable risk agent without frontier models: $0.02 per sweep, 250x cheaper than an LLM judge
alexcovo_eth · x · 2026-10-02
edwardirby shared an architecture for a reliable business-risk monitoring agent that doesn't need a frontier model, costing only $0.02 per sweep.
- Stack: You.com search + TypeSafe's Jev for typed judgments + Qwen for proposals & synthesis, integrated via MCP
- vs. a plain LLM judge: similar report quality, but the raw LLM flip-flopped on the same threat (0.35 → 0.68 → 0.50), while Jev stayed consistent
- Recall: the LLM judge missed 5 of 11 investigations; Jev missed none
- Cost/speed: 250x cheaper and 3-6x faster
Key takeaway: replacing naked LLM judging with typed, constrained outputs yields more stable and reliable agent decisions on cheap models.
More from coding & agent
- Thorsten Ball: Reconsider everything you assumed about coding models 4-6 months ago — pvncher · 2026-10-02
- Benzi: compiler-backed coding agent claims 78.2% SWE-bench Verified at under 10¢ per fix — DonkeyTheKing · 2026-10-02
- GPT-6.1 Astra builds Gaussian Splats via Blender MCP, viewed in SuperSplat — willeastcott · 2026-10-02
- Hoshikage Pulse: an Opus 5.5-built rhythm game with 128 songs, playable in browser — TAbrodi · 2026-10-02
- Conductor Mobile launches: run and control a team of cloud coding agents from iPhone — charlieholtz · 2026-10-02
- Do AI agents actually save time, or do you spend more time fixing their mistakes? — MantisReka · 2026-10-02