A new CS2 replay benchmark would ask agents to spot cheaters in impossible cases
rufreakde1 · reddit · 2026-07-23
A Reddit proposal turns CS2 cheat detection into an “impossible” benchmark for agents
The post proposes benchmarking agents on a deliberately hard task: detecting cheaters in CS2 replay files. The suggested setup has three categories:
- Easy — hvh
- Medium — 13k prem elo cheaters
- Hard — legit hacks
The idea is to keep adding more recordings and simply ask the agent to recognize cheaters in game recording files. The post also points to harbor-framework/frontier-bench and suggests that a planned “terminal bench 3” may include similar impossible benchmarks.
More from Research
- NSF and Astera Partner on Programmable Cloud Labs for AI-Ready Scientific Data — anshulkundaje · 2026-07-23
- Three classic improper integrals all collapse to √π or π — elonmusk · 2026-07-23
- Google Research studies AI agents for symptom interviews and diagnosis — gaganghotra_ · 2026-07-23
- Six practical ways to debug agent failures and keep performance from regressing — rdbms · 2026-07-23
- OSTP chief Michael Kratsios authors Genesis Mission report on AI for science — JungWooHa2 · 2026-07-23
- Hugging Face sandbox escape is being downplayed, says infosec researcher — proofreadre · 2026-07-23