Sandboxed agent scored 60% grading code reviews; shrinking the task got 12/12

Grimmoner · reddit · 2026-10-08

The author shares a sandbox design for giving an agent a real shell without trusting it, plus a failed agent-as-code-review-grader experiment and the fix.

Sandbox: dedicated Linux user without sudo; egress firewall allowing only GitHub/package registries; one-way inbox, one-way outbox with a final hash manifest, streamed back by a separate account; a plain-code checker rejects links, oversized output, and manifest mismatches. A planted prompt injection in a repo README went unexecuted and the agent flagged it.

Grading experiment (hand-built answer key, 70 items): Run 1: 41/68 (60%); forced verdicts: 44/69; bigger model: only +3. The agent never marked true claims false but let overstated claims through (73% caught on flat errors vs 45% on overstatements), quit early, and missed a DESIGN.md doc.

Fix: shrink questions to three checkable forms (line exists as quoted; number matches named list; file/hash/setting has value), allow only match/mismatch/not-found with full search paths, verify answers with plain code. Result: 12/12 on the trap set, 5/5 on a real job including a genuine catch.

Takeaways: watch which way your agent's errors lean (his failed open); shrink the question until the answer is checkable, then check with code, not another model; a real shell is fine if the box around it is boring.

Original post →

More from coding & agent

coding & agent channel →