Codex turns a joke prompt into a real paper on auditing benchmark generators
emollick · x · 2026-07-24
- A joke prompt to Codex produced a full PDF for a fake paper titled BenchBenchBench: Decision-Level Holdout Audits for Benchmark-Generation Evaluators.
- The paper is unexpectedly substantive: it proposes a holdout audit for benchmark-generating evaluators, using hidden target metrics and public scores to test whether a benchmark scorer is actually valid.
- In the reported experiment, the full evaluator reaches Spearman ρ = 0.945 and 90.8% pairwise accuracy on the composite target, but only ρ = 0.580 and 73.3% on hidden-only quality.
- The paper’s core point is that visible benchmark scores can hide large selection errors, so benchmark authors should disclose overlap between public scores and validation targets.
More from Research
- A 64-slide lecture gives a broad tour of multi-vector search models — IgorCarron · 2026-07-24
- NeurIPS 2026 workshop will spotlight failure modes of AI in biology — anshulkundaje · 2026-07-24
- S-Agent gets 46.4% on MMSI-Bench by turning spatial reasoning into action chains — 机器之心 · 2026-07-24
- If LLMs solve existence but fail universal claims, math academia may barely change — JFPuget · 2026-07-24
- A new idea compares fine-tuned weights to the base model with visualized deltas — DominiqueCAPaul · 2026-07-24
- Weekly AI issue #465 spotlights Kimi K3, Opik diagnostics, and a new paper — dl_weekly · 2026-07-24