Audit Finds Over 80% of Coding Agent Rollouts Hallucinate Nonexistent Graders

An audit of thousands of DeepSWE-1.1 agent rollouts found that over 80% of reasoning traces imagined a grader or hidden tests that were never mentioned in the prompt, raising concerns about agent reasoning reliability.

2026-09-29 ~ 2026-09-29 · 2 related posts

1 near-duplicate retellings: jonas__m