Terminal-Bench-Science 0.2 development starts, deadline Oct 5
sanmikoyejo · x · 2026-08-28
Development for Terminal-Bench-Science 0.2 has begun with a deadline of October 5. This benchmark evaluates AI agents on research workflows across scientific domains. The team is calling for practicing scientists to contribute tasks. The project is open-source on GitHub, featuring 70+ tasks spanning life, physical, earth, mathematical, and engineering sciences.
More from Research
- Qwen3.8-Next paper: matches 397B predecessor with 1/9 the training FLOPs — NielsRogge · 2026-08-28
- How 'Making It Worse' Made $20M: The Tech Behind Bodycam — aakashgupta · 2026-08-28
- NLP sarcasm detection challenge: Who is building SarcasmBench? — paul_cal · 2026-08-28
- PILOT Enables Live Self-Improvement for Long-Horizon Agents via Skill Distillation — PolyUHK · 2026-08-28
- Skild vs GEN: Both Call It In-Context Learning, Very Different Paths — gan_chuang · 2026-08-28
- Paper analyzes thermal tuning overhead in optical interconnects for MoE training — jwt0625 · 2026-08-28