vgel open-sources simple-reward-hacking: a reproducible environment for eliciting reward hacking
voooooogel · x · 2026-09-01
vgel shared the companion GitHub repo simple-reward-hacking: a code-execution environment designed to elicit reward hacking, tracked with AST-based metrics, where the reward is an easily hackable test-pass metric.
Key points:
- Dataset: a subset of SYNTHETIC-2 RL (Qwen3-32B-Hard, 1683 examples), tagged dont-train-on-this;
- Runs via a single command uv run vf-eval simple-reward-hacking, with configurable models and sampling args;
- Submitted code runs sandboxed by default via bubblewrap + prlimit.
Related event: Developer Releases Simple-Reward-Hacking Environment, Tests Gemma 3(4 posts)→
More from coding & agent
- Critique of AI Coding Benchmarks: Sparse Coverage and Low Utility — sergeykarayev · 2026-09-01
- Stop building 'AI CEOs': Production agents are much narrower — WesternVillage4581 · 2026-09-01
- Suprnova: Laravel-inspired Rust framework released, audited by AI red team — QuixiAI · 2026-09-01
- The Intent Engineering Framework: Designing Objectives and Constraints for AI Agents — PawelHuryn · 2026-09-01
- LangChain RT: Engineering becomes crucial skill alongside evals — LangChain · 2026-09-01
- Rayrun: Deploy Production-Grade MCPs in Under 1 Minute via AI — lucgagan · 2026-09-01