Measuring AI Research Ability via "Prompt Distance": Weak Models Reproduce Frontier Math Proofs
AI researcher Dan Shipper conducted an experiment to see whether weaker models could replicate the proof of the Erdős unit distance conjecture originally generated by the frontier model Astra, provided they were given conceptual hints like algebraic number theory. The findings reveal that with the right conceptual cues, less capable models can often successfully recreate breakthroughs made by frontier models, offering a fresh perspective on evaluating AI's scientific research capabilities.
已确认
- Experimental Context: Dan Shipper assigned a math challenge to models before boarding a flight, testing if they could reproduce frontier model Astra's proof of the Erdős unit distance conjecture with the help of prompts.
- Core Insight: Equipped with appropriate conceptual prompts, even weaker models are capable of reproducing discoveries made by frontier models.
为什么重要
- Dan Shipper introduced the concept of "prompt distance," suggesting that a model's intelligence and its value in scientific discovery can be quantified by observing how many hints it requires to reach a correct answer when solving newly discovered problems. This spectrum—spanning from complete uncertainty to explicit guidance—establishes a new dimension for measuring a model's research capabilities.
2026-08-03 ~ 2026-08-03 · 6 related posts
Primary sources
- [source] Can Weaker Models Replicate Frontier Discoveries with Hints? Exploring LLM Basins of Attraction — danshipper · 2026-08-03
- [source] Measuring Model Intelligence via 'Hint Distance' in Math Proofs — danshipper · 2026-08-03
- [source] Testing if GPT-4o Can Replicate Frontier Model Astra's Math Proof — danshipper · 2026-08-03
- Comparing Math Proofs: GPT-4o vs. Astra — danshipper · 2026-08-03
2 near-duplicate retellings: danshipper · danshipper