Small System 1 model Jev beats DeepSeek and grep in agent code search benchmark
bothlabs · reddit · 2026-09-29
The author built an extended CodeSearchNet Challenge benchmark (5 languages, 470 searches, 235k yes/no decisions per tool) comparing three agent code search approaches: grep with agent-written patterns, DeepSeek V4.1 Flash asked yes/no about every function, and TypeSafe's small System 1 model Jev (run via the author's open-source jevpipe CLI).
Results (F1): Jev 0.67, DeepSeek 0.64, grep 0.52. Both models clearly beat grep—e.g., for "matrix multiply", grep's patterns found none of 8 relevant functions (named mul, multiply, etc.) while Jev found all 8. Jev's edge over DeepSeek isn't a sure win (ahead in 93% of resamples), but it was more precise at every recall level, 2x faster, and slightly cheaper ($0.020 vs $0.022 per 1,000 files).
Calibration: Jev got more accurate as confidence rose (55%→98%), while DeepSeek was 0.9+ confident on 94% of answers yet only 51–59% correct below that, leaving little signal to sort by.
Two universal prompting lessons: (1) Ask about the code, not the searcher—"Does this function implement X?" beat "Is this what someone searching X wants?" by 0.03–0.11 F1. (2) Don't ask what code can't show—"read a CSV file efficiently" fails since no reader shows efficiency; ask "Does this read a CSV?" and filter for speed separately.
Caveats: one task, one judge LLM, 2019 code possibly in training data. Dataset is open-sourced on Hugging Face.
More from coding & agent
- Dagger founder: the Great CI Bottleneck of 2026 is a software problem, not hardware — msharmas · 2026-09-29
- Sonnet 5.5 lands in Factory: early tests show 'High' is the strong default — matanSF · 2026-09-29
- Full Browser Fallout Game Built With Claude Opus 5.5, Zero Texture or Sound Files — chrisfirst · 2026-09-29
- Clixad Bets on Ad-Funded AI Coding Credits Instead of $20/Month Subscriptions — Glass-Interaction972 · 2026-09-29
- Databricks: Opus 5.5 cuts coding costs 20%, GPT-6 Luna is 20x cheaper per task — pwendell · 2026-09-29
- Hindsight: Letting Your Agent Learn Without Breaking Policy — nishithreddy · 2026-09-29