Audit Finds Over 80% of Coding Agent Rollouts Hallucinate Nonexistent Graders
An audit of thousands of DeepSWE-1.1 agent rollouts found that over 80% of reasoning traces imagined a grader or hidden tests that were never mentioned in the prompt, raising concerns about agent reasoning reliability.
2026-09-29 ~ 2026-09-29 · 2 related posts
- Auditing Thousands of Rollouts: 80%+ of Coding Agents Reason About an Imagined Grader — jonas__m · 2026-09-29
1 near-duplicate retellings: jonas__m