Researcher Cross-References Model Text Outputs, Suspects Benchmark Anomalies

scaling01 · x · 2026-07-21

A developer ran a cross-entropy comparison on raw text responses from numerous models using existing benchmark data. The initial findings appear 'very suspicious,' prompting the author to publicly share the results and code for the community to inspect potential data contamination or cheating.

Original post →

More from Research

Research channel →