Overnight Test: Fable 5.1, ChatGPT Astra and Grok 4.7 All Fail at Building a Cross-Lab Agent System
adam_dorr · x · 2026-09-25
Adam Dorr gave Fable 5.1, ChatGPT Astra, and Grok 4.7 the same overnight task: using only their subscriptions (no APIs, no spending limit), build a multi-lab multi-agent swarm system sharing one bulletin board. All three failed — none could even keep their own agents' sign-ins consistent, and 30 minutes of human fixes each didn't help. He eventually got it working after hours of manual iteration with Fable 5.1 and Opus 5.5, noting the frontier remains jagged: a bulletin board shouldn't be harder than a full video game.
More from coding & agent
- Inside Quail: custom vLLM scheduler, workload-aware KV cache for 1B tok/min — sh_reya · 2026-09-25
- Anthropic's CI job volume grew 25x in six months — here's how they scaled test selection — JeremyCMorgan · 2026-09-25
- Perplexity's Fast Search powers Hermes Agent with 160ms p50 latency, free for all tiers — denisyarats · 2026-09-25
- Cursor launches Projects: one coordinator directing thousands of subagents, heavy users merge 6x more PRs — gaganghotra_ · 2026-09-25
- Paperclip lets any agent harness (Claude Code, Codex, Grok) work as employees in an AI company — AIFlow_ML · 2026-09-25
- Zapier CEO Wade Foster on grading AI fluency — and why the top rating is so rare — aakashgupta · 2026-09-25