Eval claim unravels: author admits he never inspected the outputs
xeophon · x · 2026-09-18
- Asked by paulcal what grader he used, xeophon replies he doesn't know and didn't actually look at the outputs.
- This undercuts the earlier claim that "harnesses don't matter": concluding without inspecting results becomes the point of contention itself.
More from coding & agent
- Univer ships Office Harness: isolated worktrees let agents parallel-edit connected docs — Scobleizer · 2026-09-18
- ThursdAI: TypeSafe's Jev decision model hits 70-500ms at $42/1B input tokens, outputs free — thursdai_pod · 2026-09-18
- Computer-use agents: 273 tests across 23 assistants, Muse books a haircut by phone — thursdai_pod · 2026-09-18
- AI Worth Using Podcast and OpenClaw Launch Hackathon to Build Your Startup's First AI Hire — heyneighbor · 2026-09-18
- 54k-star proxy runs Claude Code, Codex and 8 more coding agents through 50 free providers — eyishazyer · 2026-09-18
- simslim: open-source tool runs more iOS simulators on one Mac by killing unneeded daemons — tom_doerr · 2026-09-18