SmolDataEnvs ships 5.5K+ open verifiable RL tasks for training small models on real Kaggle data
SergioPaniego · x · 2026-09-28
SmolDataEnvs is out on Hugging Face: 5,500+ verifiable RL tasks for code and data science, built on real Kaggle data and led by adithyask — positioned as the next step beyond Wordle and GSM8K for hill-climbing small models.
- Deterministic grading, no LLM judge
- Fully open: RL environments, SFT trajectories, train/eval notebooks, and Harbor train/test/eval splits
- GRPO training logs for Qwen3.5 already published via Trackio
A ready-to-use recipe for teams doing RL on small models.
Related event: SmolDataEnvs Releases 5,500+ Verifiable RL Tasks on Real Kaggle Data(2 posts)→
More from coding & agent
- Agentick benchmark accepted at NeurIPS: LLM vs RL agents on same tasks, no single winner — pcastr · 2026-09-28
- Agentic commerce is still in its VHS/Betamax phase — builders are openly collaborating — jeff_weinstein · 2026-09-28
- Codex Computer Use 'Neutered' by Guardrails; Opus 5.5 Does the Job on First Try — iannuttall · 2026-09-28
- Local AI lemon inspection on a MacBook catches 4 defects out of 44 — iamrobotbear · 2026-09-28
- OpenAI co-founder Alex Atallah: a single chief-of-staff agent sacrifices your understanding everywhere — jeff_weinstein · 2026-09-28
- System 1 vs System 2 agent harnesses: bounded judgment vs open-ended planning, explained — blaizedsouza · 2026-09-28