TerminalBench-Science selection: 70 tasks chosen from 920 proposals
BenBlaiszik · x · 2026-08-28
TerminalBench-Science set a high bar for task inclusion, selecting only 70 tasks for v0.1 from 920 proposals. Each task underwent domain review, technical review, and a final difficulty calibration. The author estimates over 80 hours of effort went into each task to ensure they truly challenge top-tier models.
More from Research
- DSPy.Flex Optimizes Python Modules with Fewer LLM Calls — lateinteraction · 2026-08-28
- MIT Framework PottsMPNN Designs Proteins Beyond Natural Structures — nordicinst · 2026-08-28
- Podcast: Discussing Python Tooling Handbook and the Value of Blogs in the Age of Agents — tdhopper · 2026-08-28
- MIT's PottsMPNN Framework Advances Protein Design Beyond Natural Sequences — MIT News AI · 2026-08-28
- Debate on the Term "Double Blind" in Model Evaluation Context — BlancheMinerva · 2026-08-28
- Debugging exact DDP checkpoint resumption: CUDA atomics, Triton autotune and bucket layout pitfalls — ReinforcedKnowledge · 2026-08-28