Overnight Test: Fable 5.1, ChatGPT Astra and Grok 4.7 All Fail at Building a Cross-Lab Agent System

adam_dorr · x · 2026-09-25

Adam Dorr gave Fable 5.1, ChatGPT Astra, and Grok 4.7 the same overnight task: using only their subscriptions (no APIs, no spending limit), build a multi-lab multi-agent swarm system sharing one bulletin board. All three failed — none could even keep their own agents' sign-ins consistent, and 30 minutes of human fixes each didn't help. He eventually got it working after hours of manual iteration with Fable 5.1 and Opus 5.5, noting the frontier remains jagged: a bulletin board shouldn't be harder than a full video game.

Original post →

More from coding & agent

coding & agent channel →