Proximal details cheating-resistant task design for FrontierSWE benchmark
nrehiew_ · x · 2026-09-03
Proximal published the blog and GitHub for FrontierSWE, focusing on verifier hackability when designing genuinely difficult tasks for frontier models.
- Context: Anthropic recently detailed the link between reward-hacking and misalignment, prompting the team to share the safeguards and design decisions used to build cheating-resistant tasks.
- Author nrehiew recommends reading the blog, noting many strong contributors worked on it.
Related event: Designing Hack-Proof Benchmarks as Models Game the Verifiers(3 posts)→
More from coding & agent
- Linear's bug autofix loop closed 300+ bugs in 30 days via Datadog/Sentry and its coding agent — zeeg · 2026-09-03
- Malicious .git configs make Claude Code, Codex, Cursor run attacker code pre-trust-prompt — Thionne_WTZ · 2026-09-03
- Indie dev marclou goes agent-first: every new SaaS ships with MCP and 60+ agent tools — marclou · 2026-09-03
- Indie dev marclou goes agent-first: 60+ MCP tools make his new SaaS fully AI-operable — steipete · 2026-09-03
- 18.2M tokens of Fable 5.1 built an entire end-to-end workflow app running in the browser — gaganghotra_ · 2026-09-03
- DeepMind's 83-Page Study: Autonomous Research Agents Fabricate 90% of Findings — williamtp · 2026-09-03