DIY harness × model benchmark: Claude Code hits 100% while deepseek-v4.1-flash matches 96% in half the time

dh7net · reddit · 2026-09-28

A Reddit user built a custom benchmark spanning basic math, vision, computer use (reading emails, browsing stores) and coding to compare model × harness combinations on both capability and speed.

Top by capability

Top by speed

Key takeaway: the same model can swing tens of points across harnesses, and local small models trade hours of runtime for free inference.

Original post →

More from coding & agent

coding & agent channel →