Odysseys Benchmark for Long-Horizon Computer-Use Agents to Debut at COLM 2026
rsalakhu · x · 2026-10-09
A new benchmark called Odysseys, targeting long-horizon computer-use agents (CUAs), will be presented as poster 46 at COLM 2026 by Lawrence Jang and Jing Yu Koh.
- One of the few benchmarks specifically evaluating agents on multi-step, long-duration real computer tasks
- Details are in the poster and upcoming paper
More from Research
- Prompt Tuning Is Forgotten Lore — Are We Massively Underusing Finetuned Tokens? — cephaloform · 2026-10-09
- Ricardo Baeza-Yates Lecture: When Will ML Evaluation Stop Fooling Itself? — PolarBearby · 2026-10-09
- Alzheimer's Translation Challenge Launches With 150M-Cell Atlas and Wet-Lab Testing — xeophon · 2026-10-09
- Swapping harness lifts GPT-5.6 repo migration from 6.5% to 31%, paper finds — omarsar0 · 2026-10-09
- Microsoft's ThinkingBox: Kimi-K3 tops discovery at 93.89% but only 13.41% solve tasks 20/20 — tuhin_k · 2026-10-09
- Sakana AI's Continuous Memory Machine splits short-term compute from long-term storage — dair_ai · 2026-10-09