Developer Slams Base 1 Benchmarks: No Sample Size or Error Bars, Just a Product Update

ziv_ravid · x · 2026-08-05

Developer ziv ravid heavily criticized the recently announced benchmarks for the Base 1 model. Base 1 claimed a 74.9% success rate in building apps, closely trailing competitors, based on production A/B testing with millions of builders.

However, the critique points out that these metrics completely lack sample sizes, error bars, and clear definitions for "success" or "frustration." Because these internal tests cannot be inspected or reproduced externally, the results are dismissed as a mere product PR update rather than a rigorous scientific benchmark.

Original post →

More from Models

Models channel →