Terminal-Bench meetup draws ~150 after team once doubted finding 10 people
simonguozirui · x · 2026-09-24
Laude Institute says the Terminal-Bench team once cancelled a task-writing meetup before their first release, doubting they could get ten people in a room — tonight 150 squeezed into Laude Lab. The team covered the state of the bench and the Harbor framework, the making of Terminal-Bench-Science, and how to keep pace with better models: continuous benchmarks, real-world evals, and long-horizon multi-agent challenges.
More from Research
- Dual-H200 fine-tune of Marigold V2 with 16-frame temporal attention aims to fix video depth flicker — AntonObukhov1 · 2026-09-24
- Same Prompt, Opposite Results: GPT-4 Goes Silent 30/30 Where GPT-3.5 Never Stops — rayanpal_ · 2026-09-24
- Amazon's BoundaryMORPH uses Gaussian Processes to budget cross-encoder reranking, +5.4 nCG@100 — _reachsumit · 2026-09-24
- Paper: retrieval recall ceilings LLM recommendation reranking — oracle eval inflates NDCG up to 95%, none beat CF — _reachsumit · 2026-09-24
- Beyond a scalar: distributional serving interfaces let multiple task heads reuse watch-time distributions — _reachsumit · 2026-09-24
- Meta's Muse Realtime Avatar beats two leading commercial avatar systems in blind tests — AIatMeta · 2026-09-24