ZenML built its first RL environment to benchmark coding agents on ML workflow tasks
strickvl · x · 2026-09-11
Inspired by a viral tweet, the ZenML team built their first RL environment to test how well coding agents write ML training/deployment pipelines. Key points:
- Process largely driven by Fable with humans in the loop
- Each environment = preloaded files or a ZenML server + task instructions + a coding harness
- Multiple agents (Codex, Claude Code, Terminus 2) used to solve tasks
Author strickvl notes models appreciate structure more than expected; deeper analysis and tasks not yet saturated by frontier models are coming next week.
More from coding & agent
- Google AI Studio ships integrated docs for humans and agents — omarsar0 · 2026-09-11
- Writing great OKRs is becoming the key skill for directing AI agents — manosaie · 2026-09-11
- Spicy take making the rounds: "You need better engineers to work with LLMs" — andreisavu · 2026-09-11
- Two AI agents completed a substantial payment on their own in under five minutes — mattshumer_ · 2026-09-11
- Websites can hand tools to AI agents: WebMCP demo schedules posts from the browser — haltakov · 2026-09-11
- LangChain ships Managed Deep Agents with containerized evals via LangSmith — LangChain · 2026-09-11