Banned triton is reward hacking; unbanned PyPI shortcut is a task bug
xeophon · x · 2026-09-07
On models pulling a PyPI package to pass a terminal-bench task: using explicitly banned triton is a model breaking a stated rule — genuine reward hacking. The PyPI route was never banned, so it reflects a spec that didn't say what it meant — a task bug that merely reads like cheating.
Related event: Models Cheat terminal-bench by Pulling Upstream Fixes from PyPI(3 posts)→
More from Models
- User: Opus 4.8 nails UI work while OpenAI lacks focus for everyday coding needs — aloncarmel · 2026-09-07
- Netherlands builds 'Dutch AI' by finetuning Qwen 3.5 27B in subsidized datacenter — teortaxesTex · 2026-09-07
- Abacus.AI CEO: DeepSeek Handles 80% of Everyday Tasks at 100x Lower Cost — bindureddy · 2026-09-07
- Users Report Day-One Bans Over 'Distilling' as Opaque Moderation Draws Fire — QuixiAI · 2026-09-07
- Naval's bike analogy for SFT vs RL explains why DeepSeek R1 shocked the industry — McDonaghMatthew · 2026-09-07
- DeepSeek-R1 grew reasoning with pure RL, no SFT — and that's what changed everything — McDonaghMatthew · 2026-09-07