Split decisions from generation: 11 decision points, only 2 real generation calls per agent loop
blaizedsouza · x · 2026-09-28
- The key split in agent architecture isn't "small model vs big model" — it's deciding which calls actually need generation at all.
- A lightweight decision layer (jev) routes the next worker, scores which context is worth keeping, catches repeated actions before they burn a run, gates risky tool calls, and sends uncertainty to human review. Opus handles only the expensive parts: planning, writing, coding, synthesis.
- Rationale: a typical loop has 11 decision points but only 2 real generation calls. Separating them means latency and cost stop scaling with every tiny fork in the workflow.
- With this split, agent architecture looks surprisingly overcomplicated in hindsight.
More from coding & agent
- The Vanishing Apprentice: How AI Is Reshaping the Junior Developer Role — ArtificialOther · 2026-09-28
- Higgsfield ships 11 production skills that leave Claude with editable project files — xiaohu · 2026-09-28
- AI-generated 7-minute SQLite repo explainer stuns with coherent code walkthrough — deedydas · 2026-09-28
- Is Agentic scores how AI-agent-ready your website is, via a single npx command — seanwbren · 2026-09-28
- SolidBot moves real steel: post-processed robot programs now heading into TCP and accuracy tests — MatthewChang · 2026-09-28
- Grok Bot and Muse too dumb for business agents, says engineer comparing Claude Code — jdjohnson · 2026-09-28