9B Open-Weight Model Drops Agent Accuracy From 96% to 62.3%

TheZachMueller · x · 2026-09-30

Lambda published Open Jarvis research: dropping a 9B open-weight model (Qwen3.5-9B) into an agent system tuned for a frontier cloud model (Claude Opus 4.6) cut accuracy on PinchBench — which measures how well an LLM performs as the brain of OpenClaw agents — from 96.0% to 62.3%.

The bottleneck is the harness: the configuration that tells the model what to do, which tools it can use, and how to learn from mistakes. Retargeting the spec around the local model — changing quantization, the reasoning loop, tool access, and memory — recovered accuracy to 88.4%, closing 77% of the gap through harness tuning alone, before any further spec-search optimization.

Open Jarvis is built around a spec unique to a model and the machine it runs on, with five independently configurable primitives (Intelligence, Engine, Agent, Tools & Memory). An LLM-guided spec search uses a frontier cloud model to diagnose failures and propose edits. The work is open source and evaluated on Lambda GPUs.

Original post →

More from coding & agent

coding & agent channel →