Yarin Gal: LLM experiments that don't replicate are just failures, and an AI arXiv could help
yaringal · x · 2026-09-23
In a follow-up to his call for arXiv to reject LLM-written papers, Oxford's Yarin Gal argues that non-replicating LLM-written experiments should be treated like any other failed experiment, and suggests an 'AI arXiv' for LLM-written papers. He maintains that filtering only LLM-written prose (not LLM-assisted experiments/proofs) would drastically cut slop by raising the cost of spamming.
Related event: Oxford's Yarin Gal: Irreproducible LLM Experiments Are Just Bad Experiments(2 posts)→
More from Research
- arXiv math submissions jump 33.5% in 2026, with 29 of 30 subfields growing — rohanpaul_ai · 2026-09-23
- Mixture-of-Depths routes FLOPs per token to cut transformer compute — burny_tech · 2026-09-23
- AI has solved the hardest part of formal verification, says Theorem co-founder — burny_tech · 2026-09-23
- Minerva framework preprint is out, says Brian Hie — BrianHie · 2026-09-23
- Genome language models uncover new class of reverse-transcriptase mechanisms — BrianHie · 2026-09-23
- Mathematician shares a cheap 4-step heuristic for hyperparameter tuning — dejanseo · 2026-09-23