Paperena Benchmarks AI Scientists Across Full Research Cycles: Writing, Reviewing, Revising

yeewhye · x · 2026-09-24

SnorkelAI presented Paperena at the Frontier Data Summit: a benchmark that evaluates AI scientists over realistic research cycles—writing, reviewing, and revising papers—scoring factuality, reproducibility, novelty, and community behavior. The invite-only October 2026 summit also features 25+ new benchmarks (STELLA-Bench, CollusionBench, Terminal Bench 2.1) and speakers including François Chollet and Dawn Song.

Original post →

More from Research

Research channel →