15 harnesses, same result: exchange over whether eval harness choice even matters
paul_cal · x · 2026-09-18
- xeophon reports running Hello World with Fable 5.1 at max settings across 15 harnesses, getting identical results, and concludes harnesses don't matter.
- paulcal presses on what grader was used; xeophon admits he didn't check and never actually inspected the outputs.
- The exchange raises doubts about whether sweeping conclusions on eval harnesses are credible without output inspection.
Related event: 15 harnesses, identical results spark eval benchmark debate(2 posts)→
More from coding & agent
- Univer ships Office Harness: isolated worktrees let agents parallel-edit connected docs — Scobleizer · 2026-09-18
- ThursdAI: TypeSafe's Jev decision model hits 70-500ms at $42/1B input tokens, outputs free — thursdai_pod · 2026-09-18
- Computer-use agents: 273 tests across 23 assistants, Muse books a haircut by phone — thursdai_pod · 2026-09-18
- AI Worth Using Podcast and OpenClaw Launch Hackathon to Build Your Startup's First AI Hire — heyneighbor · 2026-09-18
- 54k-star proxy runs Claude Code, Codex and 8 more coding agents through 50 free providers — eyishazyer · 2026-09-18
- AI Guard maintainer: commenter found flaw in our init-context measurement plan — Yashhh_21 · 2026-09-18