Hiding Y-Axis Labels in Early nanogpt Benchmarks Is "Academic Dishonesty"
PMinervini · x · 2026-10-06
A debate over benchmark integrity in open model releases: critics called it "academic dishonesty" to hide y-axis labels while benchmarking against the first 3% of a nanogpt training run. The quoted post explains the substance — training only until 5.0 nats of fineweb validation loss is so high it doesn't even appear on the nanogpt-small training graph; at that stage most of the signal is still just pushing logits of unlikely tokens to negative infinity, and real feature learning hasn't begun. Early-training snapshots therefore make misleading comparison charts.
More from Models
- Grok web app to add Team Bots: one shared bot, private chats for each member — nima_owji · 2026-10-06
- Liquid AI's d1 hits 85-97% on VisA inspection tasks with zero task-specific training — JosephJacks_ · 2026-10-06
- QuantCode: domain pretraining + SFT lifts Qwen trading-code pass from 27.8% to 58.2% — Alexey Chernysh · 2026-10-06
- GLM-5.3 appears context-aware, reportedly saying it's running low on context — khademinori · 2026-10-06
- Azure catalog briefly leaks unreleased OpenAI model GPT-6.1-Sol, page pulled within an hour — _AndrewZhao · 2026-10-06
- Opus 5.5 is efficient on subscription, not via API — 6.1 remains the workhorse — haider1 · 2026-10-06