Berkeley team unveils DeepScholar-Bench, a live benchmark for AI research synthesis
berkeley_ai · x · 2026-10-06
A UC Berkeley team including Liana Patel, Ion Stoica, Matei Zaharia and Carlos Guestrin is presenting DeepScholar-Bench at COLM 2026, a live benchmark and automated evaluation framework for generative research synthesis systems.
- Task: systems retrieve from the live web, synthesize and cite prior work to generate a full related-work section — far beyond short factual QA.
- Data: queries and human-written exemplars come from recent high-quality ArXiv papers, sidestepping staleness and contamination in curated datasets.
- Evaluation: holistic scoring across knowledge synthesis, retrieval quality, and verifiability.
- Bonus: an open-source reference pipeline, DeepScholar-ref, built on the LOTUS framework as a strong baseline.
The author will also present a Multi-Agent Transactive Memory paper at the Lifelong Agents workshop.
More from coding & agent
- Open-source LCU decouples Codex Computer Use so any AI agent can drive it via MCP — xiaohu · 2026-10-06
- Indie dev's agents auto-produced 53 pregnancy reels at ~$0.03 GPU cost per clip, skill released free — victor_explore · 2026-10-06
- Jev playbook: use a $0.042/M-token judge model to route agents and clean context — blaizedsouza · 2026-10-06
- "I don't code anymore, I just yap and stuff happens" — jaivinwylde · 2026-10-06
- Prompt craft: using Opus 5.5 with JS Canvas to draw a Chinese seal-script seal — dotey · 2026-10-06
- "It's insane that coding is now just audibly speaking at your computer" — KevinNaughtonJr · 2026-10-06