9B Open-Weight Model Drops Agent Accuracy From 96% to 62.3%
TheZachMueller · x · 2026-09-30
Lambda published Open Jarvis research: dropping a 9B open-weight model (Qwen3.5-9B) into an agent system tuned for a frontier cloud model (Claude Opus 4.6) cut accuracy on PinchBench — which measures how well an LLM performs as the brain of OpenClaw agents — from 96.0% to 62.3%.
The bottleneck is the harness: the configuration that tells the model what to do, which tools it can use, and how to learn from mistakes. Retargeting the spec around the local model — changing quantization, the reasoning loop, tool access, and memory — recovered accuracy to 88.4%, closing 77% of the gap through harness tuning alone, before any further spec-search optimization.
Open Jarvis is built around a spec unique to a model and the machine it runs on, with five independently configurable primitives (Intelligence, Engine, Agent, Tools & Memory). An LLM-guided spec search uses a frontier cloud model to diagnose failures and propose edits. The work is open source and evaluated on Lambda GPUs.
More from coding & agent
- OpenRoboto runs open robot intelligence contests on Bittensor, miners evolve shared base models — markjeffrey · 2026-10-01
- Weco agent rewrote its own scoring code; founder says lock eval files before agent runs — victor_explore · 2026-10-01
- Real-world agent finance ops: escalates EUR 9,000 bill over approval limit, avoids duplicate payments — kimmonismus · 2026-10-01
- Agent value isn't measured in minutes saved, but in background execution and trust — itsOmSarraf_ · 2026-10-01
- Agent Verification Startup Beltic Emerges From Stealth With $8.8M to Make Agent Traffic Verifiable — alifcoder · 2026-10-01
- OpenAI marketplace may not require OpenAI inference, hinting at harness-level competition — matt_slotnick · 2026-10-01