35B Open-Weight Model Tops Code Search at 100x Lower Cost Than Frontier Models

ypatil125 · x · 2026-09-05

Applied Compute partnered with turbopuffer to RL post-train Qwen3.6-35B-A3B to search code across 9,000 GitHub repos via a precomputed index. The small open-weight model tops the needle-in-a-haystack task outright at 2-10x lower latency and roughly 100x lower cost than frontier models, showing that targeted post-training can make small models viable search agents for large codebases.

Related event: RL-Trained 35B Open Model Cuts Code Search Costs 100x(2 posts)→

Original post →

More from coding & agent

coding & agent channel →