Prompt Distance: Can Weaker Models Reproduce Frontier Proofs?
AI researcher Dan Shipper conducted an experiment testing whether weaker models can reproduce a frontier model's proof of the Erdős planar unit distance conjecture when given conceptual prompts such as algebraic number theory. The experiment showed that with appropriate conceptual prompts, weaker models often reproduce frontier findings, offering a new perspective on evaluating AI research capability.
Confirmed
- Experiment background: Dan Shipper set a math challenge before boarding, testing whether models could reproduce Astra's proof of the Erdős planar unit distance conjecture with prompts.
- Core claim: With appropriate conceptual prompts, weaker models can often reproduce frontier findings.
- Comparison: The developer shared GPT-4o's analysis results and directly compared them with OpenAI's frontier model Astra's solution, further validating whether weaker models can match frontier reasoning with proper prompts.
Why it matters
- Dan Shipper proposes the concept of 'prompt distance'—measuring how many prompts a model needs to reach the correct answer for a new discovery—as a metric to quantify AI intelligence and research value. This span from complete ignorance to explicit guidance offers a new dimension for assessing model research capability.
2026-08-03 ~ 2026-08-03 · 6 related posts
- Episode 1: OpenAI Tests Multi-Agent Model Astra, Demoed to US Officials(2026-08-01, 17 posts)
- Episode 2: OpenAI's Internal Model Astra Cracks 10 Major Math Problems(2026-08-01, 99 posts)
- Episode 3: Google's Astra Solves 10 Scientific Problems for Under $2,000(2026-08-01, 4 posts)
- Episode 4: OpenAI Math Breakthrough Questioned by Gary Marcus and Others(2026-08-01, 38 posts)
- Episode 5: Rumor: OpenAI to Release Astra Model Focused on Multi-Agent Collaboration(2026-08-02, 3 posts)
- Episode 6: Anthropic Employee Reproduces Half of Astra Math Proofs Using Fable in 24 Hours(2026-08-02, 4 posts)
- Episode 7: Prompt Distance: Can Weaker Models Reproduce Frontier Proofs?(2026-08-03, 6 posts)
- Episode 8: OpenAI Claims Internal Model Solves 10 Math Problems for $2000, Sparking Debate(2026-08-03, 15 posts)
- Episode 9: OpenAI Releases Astra Mathematical Proofs and Reasoning Manuscripts(2026-08-04, 4 posts)
- Episode 10: OpenAI Rumored to Release Astra Next Week, Possibly Largest Pretrained Model Since GPT-4.5(2026-08-06, 10 posts)
- Episode 11: OpenAI Slows Astra Development Over Cyber Risk, Limits Initial Release(2026-08-08, 30 posts)
- Episode 12: OpenAI Teases Astra as Its First Critical Cybersecurity Model(2026-08-08, 4 posts)
- Episode 13: OpenAI's Astra Model May Launch in August with 10T Parameters(2026-08-09, 6 posts)
- Episode 14: OpenAI's Next Model Codenamed Doug, Largest Pretraining Run Planned by Year-End(2026-08-09, 8 posts)
- Episode 15: OpenAI Pauses Astra Model Development Over Cybersecurity Risks(2026-08-10, 3 posts)
- Episode 16: OpenAI Faces Internal Dispute Over AI Model Safety Ratings(2026-08-11, 2 posts)
- Episode 17: Polymarket Odds Put 76% Chance on OpenAI's Astra Launching Next Month(2026-08-14, 2 posts)
- Episode 18: OpenAI's Next Model 'Astra' Rumors Heat Up as Employees Tease Launch(2026-08-15, 5 posts)
- Episode 19: Power Bottlenecks Threaten US AI Data Center Expansion(2026-08-17, 4 posts)
- Episode 20: OpenAI Halts Frontier RL Training as Astra Hits Critical Cyber Threshold(2026-08-17, 17 posts)
Primary sources
- [source] Can Weaker Models Replicate Frontier Discoveries with Hints? Exploring LLM Basins of Attraction — danshipper · 2026-08-03
- [source] Measuring Model Intelligence via 'Hint Distance' in Math Proofs — danshipper · 2026-08-03
- Testing if GPT-4o Can Replicate Frontier Model Astra's Math Proof — danshipper · 2026-08-03
- [source] Comparing Math Proofs: GPT-4o vs. Astra — danshipper · 2026-08-03
2 near-duplicate retellings: danshipper · danshipper