Ethan Mollick: AI research on older models requires caution
emollick · x · 2026-08-24
Ethan Mollick discusses the pitfalls of publishing AI impact research based on older models like GPT-4. He argues that positive claims (e.g., AI crossed a threshold) are generally more durable than negative ones.
If the finding is that AI CAN do something:
Generally valid, as abilities rarely regress.
If the finding is that AI CANNOT do something (Requires Caution):
- Be Explicit: State "GPT-4 could not do X" to serve as a benchmark rather than a general claim.
- Measure Trends: Compare across model generations (including reasoning models) to track progress.
- Argue Inherent Flaws: Provide strong evidence for a natural limitation that prevents AI from doing X.
- Focus on Moderators: Study how prompts, context, or social factors affect performance, which remains relevant.
- Focus on Humans: Examine human reactions and social impacts.
More from Research
- Jitendra Malik: Stop Conflating VLMs with World Models in Robotics — JitendraMalikCV · 2026-08-24
- Six months of evals: never let the model rewrite the source; BM25 weighting lifts recall@10 to 0.86 — Cryvixx · 2026-08-24
- Japanese Team Creates Female Clones from Male Mouse Cells Using CRISPR — Promptmethus · 2026-08-24
- NVIDIA AVO Scores 100% on ARC-AGI-3, Proving System Design Trumps Model Capability — cantrell · 2026-08-24
- Continual learning should adapt to noisy data, not over-clean it — Shahules786 · 2026-08-24
- IR Papers Vol.170: RAG Effectiveness & Agent Retrieval — _reachsumit · 2026-08-24