Optimized open-source Argus hits SOTA on WGO-Bench, lifting semantic F1 from 29.8% to 50.2%
_sonith · x · 2026-10-06
MidcenturyAI spent a weekend optimizing the open-source video annotation model Argus and pushed it to SOTA on WGO-Bench: semantic F1 jumped from the 29.8% out-of-the-box score to 50.2%, well ahead of other leading annotation services.
All gains came from how the model sees video:
- Denser frame inputs
- A second pass on each action's boundaries
- A final consistency check on every label
The team credits Pantheon and the community for open-sourcing Argus. The full repo is coming soon; a tech report is available now. A nice case study of community optimization quickly surpassing an open-source baseline.
More from Research
- Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test — alexisjross · 2026-10-06
- Embedding Every Font with Neural Networks Yields a Flower-Shaped Map of Google Fonts — Chroma-Crash · 2026-10-06
- Crawler Zoo Launches a Free Arena for Testing Local-Model Agents — Time_Instruction_955 · 2026-10-06
- Trained agentic context management: 8K-context small model matches GPT-5.4 at 1M on OOLONG — xennygrimmato_ · 2026-10-06
- User Sim Index is broken: trivial bot scores 95% across behavioral dims — ericzelikman · 2026-10-06
- Used OpenAI Dots as a Free Agent Swarm to Break a 47-Year-Old Math Record — jaxchang · 2026-10-06