ByteDance Releases StartupBench: Top Models Complete Only 30% of Tasks

ByteDance-Seed · hf · 2026-08-19

ByteDance released StartupBench to evaluate general-purpose agents on real-world startup workflows. It reveals that even top models complete only about 30% of tasks, highlighting gaps in instruction following and domain expertise.

Original post →

More from coding & agent

coding & agent channel →