Blind Test Pits AI Models Against Real Community and Social Media Tasks
thisiskp_ · x · 2026-10-08
thisiskp previewed a multi-model blind comparison: models will tackle real tasks from two jobs — community management and social media. Tania picks community asks from a menu; social tasks are chosen by the audience in chat.
Every model gets the same prompt via Netlify AI Gateway, outputs are shown anonymously as A, B, C, and D, and the audience votes for a favorite before the models are revealed.
More from coding & agent
- Agents love tidying up files nobody asked them to touch — JFPuget · 2026-10-08
- Open-source bridge turns a browser chat tab into an OpenAI-compatible API endpoint — harshanacz · 2026-10-08
- GraphRAG Bug: Deleted Documents Stay Indexed and Retrievable After Updates — JeremyCMorgan · 2026-10-08
- Paper: Vibe Coding Kills Open Source as AI-Recommended Repos Lose Stars — soumitrashukla9 · 2026-10-08
- Every publishes definitive guide to Compound Engineering, the philosophy behind 7k-star plugin — every · 2026-10-08
- Claude Haiku 5.5 matches GPT-6 Luna pricing but a stingier tokenizer hides a 1.25x cost hike — Simon Willison · 2026-10-08