Running the same Dots dashboard prompt 8x: 12 providers, no stack specified

edwin · x · 2026-10-03

edwin ran a controlled experiment: ask OpenAI's Dots to build an AI status dashboard aggregating API health for 12 providers including Claude, Cursor, OpenRouter, Gemini, Mistral, and Cohere — public page, incident history, private watchlists, 15-minute refresh — without specifying any stack, then repeated the identical experiment 8 times to observe variance in the agent's choices and delivery.

Follow-up posts reveal uneven results on scraping and scheduling: some builds shipped with no active schedule, and hand-rolled parsers replaced off-the-shelf tools. A rare 'same prompt 8 times' look at agent reliability.

Related event: Developer runs OpenAI Dots 8 times to build AI status dashboard(2 posts)→

Original post →

More from coding & agent

coding & agent channel →