Small System 1 model Jev beats DeepSeek and grep in agent code search benchmark

bothlabs · reddit · 2026-09-29

The author built an extended CodeSearchNet Challenge benchmark (5 languages, 470 searches, 235k yes/no decisions per tool) comparing three agent code search approaches: grep with agent-written patterns, DeepSeek V4.1 Flash asked yes/no about every function, and TypeSafe's small System 1 model Jev (run via the author's open-source jevpipe CLI).

Results (F1): Jev 0.67, DeepSeek 0.64, grep 0.52. Both models clearly beat grep—e.g., for "matrix multiply", grep's patterns found none of 8 relevant functions (named mul, multiply, etc.) while Jev found all 8. Jev's edge over DeepSeek isn't a sure win (ahead in 93% of resamples), but it was more precise at every recall level, 2x faster, and slightly cheaper ($0.020 vs $0.022 per 1,000 files).

Calibration: Jev got more accurate as confidence rose (55%→98%), while DeepSeek was 0.9+ confident on 94% of answers yet only 51–59% correct below that, leaving little signal to sort by.

Two universal prompting lessons: (1) Ask about the code, not the searcher—"Does this function implement X?" beat "Is this what someone searching X wants?" by 0.03–0.11 F1. (2) Don't ask what code can't show—"read a CSV file efficiently" fails since no reader shows efficiency; ask "Does this read a CSV?" and filter for speed separately.

Caveats: one task, one judge LLM, 2019 code possibly in training data. Dataset is open-sourced on Hugging Face.

Original post →

More from coding & agent

coding & agent channel →