VHD-Play Generates 3,300 Agentic RL Environments at a Few Cents Each
Xinjie Shen · hf · 2026-09-24
VHD-Play reverses typical environment-generation pipelines: it first samples and solves a mathematical model, then a corpus-grounded setter renders the solved mechanism's decision process as stateful tools, with executable dynamics and trajectory-scoring reference inherited from the same model—avoiding post hoc alignment between environments and outcome rules.
- Produces 3,300 diverse agentic environments at a few cents each
- Training Qwen3.6-35B-A3B on three families raises mean agentic score from 0.204 to 0.815 on a five-family diagnostic
- Gains generalize to held-out instances, eight unseen mechanism families, and external benchmarks for function calling, travel planning, and e-commerce
- On E-Commerce Bench the trained checkpoint never goes bankrupt and exceeds Qwen3.7-Max
- Most of the learnable gap lies in stateful interaction, not underlying problem solving
More from coding & agent
- A Hermes Agent prompt that audits Claude's memory files every session — alexcovo_eth · 2026-09-24
- Open-Source MCP Server Connects Claude/Cursor to ComfyUI with 50+ Tools — Less_Actuary_9441 · 2026-09-24
- Using Notion as a data hub makes switching AI services painless — ivanhzhao · 2026-09-24
- Hot take: a 22-year-old fluent in coding agents beats a lazy senior dev — but watch out for 'slop grenades' — jobergum · 2026-09-24
- Blender MCP hits 29k stars as Opus 5.5 builds cities on medium effort — sidahuj · 2026-09-24
- AI Agent Swarm Reverse-Engineers 2001 GBA Game Snood Byte-for-Byte in Two Weeks — Aizkmusic · 2026-09-24