Paperena Benchmarks AI Scientists Across Full Research Cycles: Writing, Reviewing, Revising
yeewhye · x · 2026-09-24
SnorkelAI presented Paperena at the Frontier Data Summit: a benchmark that evaluates AI scientists over realistic research cycles—writing, reviewing, and revising papers—scoring factuality, reproducibility, novelty, and community behavior. The invite-only October 2026 summit also features 25+ new benchmarks (STELLA-Bench, CollusionBench, Terminal Bench 2.1) and speakers including François Chollet and Dawn Song.
More from Research
- Claude discovers unknown enzyme system in phage DNA, resembling CRISPR — jarrodwatts · 2026-09-24
- Training real hardware, not a digital twin: learning directly in nonlinear wave systems — bravo_abad · 2026-09-24
- Anthropic says Claude discovered an unknown enzyme system hidden in phage DNA — nptacek · 2026-09-24
- Anthropic wet lab: ~1,000 Claude agents find unknown enzyme system in 21 hours — HealthcareAIGuy · 2026-09-24
- 950 agents ran for 21 hours and the enzyme's function is still unresolved — ns123abc · 2026-09-24
- ICLR 2027 opens reviewer bidding before NeurIPS 2026 results, researchers object — rao2z · 2026-09-24