OSWorld-Science Debuts: 146 Tasks Test How Well VLM Agents Handle Scientific Software
SciAILab · hf · 2026-10-01
- OSWorld-Science is a benchmark and evaluation environment for computer-use agents built on VLMs, covering scientific workflows like molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation.
- It includes 146 high-quality tasks designed via expert proposals and human–AI co-design, with execution-based evaluators that inspect artifacts (molecular structures, segmentation masks, plots, numerical results) and award partial credit.
- The team evaluated 12 VLMs using a harness with model adapters, interaction-loop control, and trajectory logging.
- Findings: even state-of-the-art VLMs with a strong harness struggle on key scientific tasks; the paper analyzes effects of multilinguality, reasoning effort, and context length, and outlines directions for future development.
More from Research
- LANTERN uses LLM internal activations to surface four novel OEIS integer sequence relations in under 8 hours — Pavel Tikhonov · 2026-10-01
- The geometry of inference in transformer residual streams: how predictions sharpen with depth — Timur Mudarisov · 2026-10-01
- How many context tokens does a language model actually use? Measuring effective attention set size — Timur Mudarisov · 2026-10-01
- Kernelised functional Bregman divergences paper accepted at NeurIPS 2026 NeurReps workshop — FrnkNlsn · 2026-10-01
- Why vision and hearing dominate experience: bigger cortical regions, not a saliency system — Sauers_ · 2026-10-01
- Hidden Dates in System Prompts Swing LLM Eval Scores by Up to 14% — Mario Sanz-Guerrero · 2026-10-01