TermGrade ships 1,004 graded terminal environments and 36k RL trajectories, fully open
huggingface · x · 2026-10-09
ai& Research released TermGrade on Hugging Face: 1,004 terminal environments for RL, each a real Linux task with pytest verification, plus all 36,144 graded trajectories (14,234 failures included) from six models including Kimi-K3, DeepSeek-V4-Pro, GLM-5.2, Qwen3.6-27B and gemma-4-31b-it.
- Difficulty is model-dependent: one model's 50%-pass tasks are mostly a different set for another, so grade against the model you train
- Training gemma-4-31b-it on its half-solved tasks gained +3.1 on Terminal-Bench 2.1; 5 repeat runs all beat base
- Backed by 20k GPU-hours and 4.5B agent rollout tokens; environments, trajectories, training split and checkpoint fully open
Related event: TermGrade open-sources 1,000+ executable terminal environments for agent RL(2 posts)→
More from coding & agent
- Stop wasting Opus: a 5-layer Claude Code setup with Sonnet, Haiku and Opus on call — Arindam_1729 · 2026-10-09
- $200 Max plan greentext: agents burn 29% quota, touch zero files, run no tests — BLUECOW009 · 2026-10-09
- Claude Code 2.1.295 Adds Blocking Hooks and Terminal Status Support, Prompt Tokens Up 23.5% — ClaudeCodeLog · 2026-10-09
- Claude Code 2.1.295 ships 143 CLI changes, adds blocking hooks and stream timeout controls — ClaudeCodeLog · 2026-10-09
- Claude Code 2.1.295 ships 143 CLI changes: failing hooks now block actions, gateway TTFB timeout added — ClaudeCodeLog · 2026-10-09
- Senzing ships Claude Code plugin bringing entity resolution to coding agents — JeremyCMorgan · 2026-10-09