Burning Codex Quota on a Multi-Model Agent Workflow, Then Hitting Thread Limits
craigbalding · x · 2026-09-16
The author ran a multi-agent, multi-model Codex workflow (ChatGPT Pro quota down to 14% after two resets) to transliterate Safeyolo's AI security proxy from Python to Rust. His model split: Astra xhigh as coordinator, Luna xhigh for coding, Astra low for dedicated verification, with explicit Luna xhigh overrides on two coding workers. The harness then hit agent-thread limits three times, forcing silent reuse of a Luna worker after source freeze and killing the independent reviewer — leaving expensive Astra xhigh doing 'narrow' checks. Lesson: don't silently reuse older higher-cost workers, but thread limits make cost-optimized agent splits fragile.
Related event: ChatGPT Pro User Burns Through 20x Quota After Two Resets(2 posts)→
More from coding & agent
- ModularRSI: Modular, benchmark-disjoint framework for generalizable agent harness self-improvement — IQuestLab · 2026-09-16
- Developer ditches Astra for 5.6 Sol high/xHigh setup with Luna subagent, says usage improved — rudrank · 2026-09-16
- Codex power users stuck: 7% quota left, 3-day wait, 20x plan paused — jasonkneen · 2026-09-16
- Open-source Orca runs 5 Claude Code agents in parallel, hits 60k GitHub stars — alex_verem · 2026-09-16
- claude-reflect: open-source tool turns your corrections into permanent memory for Claude Code — tom_doerr · 2026-09-16
- Free Complete Guide to Obsidian Automation released, covering AI agents on a 20,000-note vault — dSebastien · 2026-09-16