Dev claims 'Sonnet 5.5' hits 120-130tps with 100% pass rate on his end-to-end app benchmark

julianharris · x · 2026-10-01

Developer julianharris reports that an unreleased "Sonnet 5.5" is sustaining 120-130 tokens/s and passing 100% of tasks so far on his self-built spec-driven long-form benchmark, which has LLMs build complete apps end to end and tests them against hidden holdout tests. The model has not been officially announced, so the name and scores remain unverified. The linked GitHub repo (awesome-local-ai) collects install scripts for running local models like Qwen on 24GB/64GB hardware via llama.cpp, sglang and mtplx.

Original post →

More from coding & agent

coding & agent channel →