Agentick benchmark accepted at NeurIPS: LLM vs RL agents on same tasks, no single winner
pcastr · x · 2026-09-28
Agentick, a unified benchmark for training and evaluating general sequential decision-making agents, has been accepted to NeurIPS E&D.
- RL agents, zero-shot LLMs, VLMs, hybrids, bots, and humans are all evaluated on the same tasks, seeds, and score.
- The goal is to put zero-shot LLM agents and neural nets trained from scratch with RL on common ground to study their respective strengths and limitations.
- First result: no single agent dominates.
- Paper, code, and blog are available.
More from coding & agent
- Scheduled Tasks for CLI Coding Agents via herdr + Custom /routine Skill + OpenCode — swaroopch · 2026-09-29
- Manus Cue Agents Run as Independent Entities With Own Number, Inbox and Card — ChrisUniverse · 2026-09-29
- Dev: 'Showing your prompt' is now meaningless — context is references, skills and examples — trq212 · 2026-09-29
- Dev rebuilds childhood arcade game Off-Road with Codex and GPT-6, real spinning wheel controller — CSProfKGD · 2026-09-29
- AI voice agents now settle medical bills in one call — voice is the easy part — alex_verem · 2026-09-29
- DIY scheduled tasks for CLI coding agents with OpenCode, herdr and a custom skill — swaroopch · 2026-09-29