Hiding Y-Axis Labels in Early nanogpt Benchmarks Is "Academic Dishonesty"

PMinervini · x · 2026-10-06

A debate over benchmark integrity in open model releases: critics called it "academic dishonesty" to hide y-axis labels while benchmarking against the first 3% of a nanogpt training run. The quoted post explains the substance — training only until 5.0 nats of fineweb validation loss is so high it doesn't even appear on the nanogpt-small training graph; at that stage most of the signal is still just pushing logits of unlikely tokens to negative infinity, and real feature learning hasn't begun. Early-training snapshots therefore make misleading comparison charts.

Original post →

More from Models

Models channel →