Dev claims 'Sonnet 5.5' hits 120-130tps with 100% pass rate on his end-to-end app benchmark
julianharris · x · 2026-10-01
Developer julianharris reports that an unreleased "Sonnet 5.5" is sustaining 120-130 tokens/s and passing 100% of tasks so far on his self-built spec-driven long-form benchmark, which has LLMs build complete apps end to end and tests them against hidden holdout tests. The model has not been officially announced, so the name and scores remain unverified. The linked GitHub repo (awesome-local-ai) collects install scripts for running local models like Qwen on 24GB/64GB hardware via llama.cpp, sglang and mtplx.
More from coding & agent
- Two-Person Team Asks: What's the Best Production LLM Router After a Provider Outage? — DuanesKrasner · 2026-10-01
- Incident Arena benchmark puts coding agents on call for incident response — amankhan · 2026-10-01
- Google: Gemini 4 Argon agents migrate 800k+ lines of C/C++ to Rust, free 300 TiB — xennygrimmato_ · 2026-10-01
- Edge Python: A 200KB Rust/WASM Sandbox for Running Untrusted Agent Code — Healthy_Ship4930 · 2026-10-01
- Dev Open-Sources Queue Workbench to Fix ComfyUI's Nearly Unmanageable Queue — Smegnigma · 2026-10-01
- Making a Paper-Art Game Intro in Godot with Claude Opus: Full Workflow — Equivalent-Future439 · 2026-10-01