Developer Slams Base 1 Benchmarks: No Sample Size or Error Bars, Just a Product Update
ziv_ravid · x · 2026-08-05
Developer ziv ravid heavily criticized the recently announced benchmarks for the Base 1 model. Base 1 claimed a 74.9% success rate in building apps, closely trailing competitors, based on production A/B testing with millions of builders.
However, the critique points out that these metrics completely lack sample sizes, error bars, and clear definitions for "success" or "frustration." Because these internal tests cannot be inspected or reproduced externally, the results are dismissed as a mere product PR update rather than a rigorous scientific benchmark.
More from Models
- Best Local LLMs for Coding on a 128GB Mac? — Electronic_Back1502 · 2026-08-05
- Open Source AI Hits a Wall in Long-Running Agentic Loops — bindureddy · 2026-08-05
- Dev: I'd rather iterate 10 times with Gemini 3.6 Flash than wait 2 hours with Qwen 3.8 max — DynamicWebPaige · 2026-08-05
- Investor Hints at Upcoming FLUX 3 Model with Video Prompt — venturetwins · 2026-08-05
- From People-Pleaser to Stubborn: Users Complain About New Model Behavior — Hatrct · 2026-08-05
- Why Frontier LLMs Act Like Jerks: Blame RLVR Training — SpiritRealistic8174 · 2026-08-05