Snorkel AI lands five NeurIPS papers; hardest long-horizon agent tasks see <1% pass rate
ajratner · x · 2026-10-02
Snorkel AI's research team has five papers accepted to NeurIPS 2026, focused on agent evaluation and engineering:
- Agents' Last Exam: long-horizon professional tasks; <1% average pass rate on the hardest tier.
- Continual Learning Bench: do agents actually improve with experience?
- JudgmentBench: expert attorney preference judgments for ranking legal AI output.
- SkillOrchestra: routing tasks by agent competence and cost.
- SlopCodeBench: agent-written code gets more bloated and eroded with every iteration.
Together they quantify how far agents remain from reliable long-horizon professional work, with striking data points like sub-1% pass rates and per-iteration code degradation.
More from coding & agent
- SerenityOS Creator Maps Open Source's Five Stages of Grief Over AI Code — vivekhaldar · 2026-10-02
- Solo Founder Launches OpenSwitchboard, a Matchmaker for AI Assistants Handling Real-World Errands — EnvironmentalRice348 · 2026-10-02
- App idea: an MCP-powered alarm that wakes you when your agent finishes — Angaisb_ · 2026-10-02
- Wellness Project Launches MCP Connector Bringing Health Data Into Claude — turnnoblindeye · 2026-10-02
- How an FDE maps business processes for AI deployment using three sources — vasuman · 2026-10-02
- Rhyven Marketplace: Open-Source MCP-Based App Store for AI Agents Launches Linux Preview — Due-Telephone8276 · 2026-10-02