AI-slop blind test data and reproducible scripts land on GitHub
PawelHuryn · x · 2026-09-07
Pawel Huryn published the AI-slop blind test's raw data, scripts and results to his GitHub experiments repo. Everything is reproducible: exact scripts, unedited logs, and READMEs with method, sample size and caveats; the repo also hosts prior experiment sets like bug-hunt-bench, five-models-three-harnesses, and stealth system-prompt fingerprinting. He asks if followers want it as a live benchmark like Bug Hunt Bench.
Related event: Blind Test of 12 Models: Gemini 3.8 Most 'AI-Flavored'(3 posts)→
More from Models
- Leak hints a 20T-parameter model is on the way — Dr_Singularity · 2026-09-07
- Blogger revises AI training scale estimate to 5-7T, says 8T already too generous — scaling01 · 2026-09-07
- OpenAI's Astra math 'breakthroughs' commit research misconduct, mathematicians say — asusarla · 2026-09-07
- Peter Gostev debunks model sparsity leak: Kimi 26:1, DeepSeek 32:1, 1.2T active params implausible — inductionheads · 2026-09-07
- OpenAI's newest models block function tools on /v1/chat/completions, forcing Responses API migration — AI-Specialist-6597 · 2026-09-07
- Jensen Huang declares "AGI is here"; Chinese LLMs top global inference usage for 19th week — 快鲤鱼 · 2026-09-07