ExecRetrieval benchmark shows code embedders rank buggy near-clones above correct code
kennesaw · hf · 2026-09-03
ExecRetrieval tests whether code embeddings can distinguish correct implementations from near-identical buggy variants, finding that leading retrievers frequently rank incorrect near-clones above canonical solutions.
More from Research
- LeVJEPA: video encoder matches V-JEPA 2 with 5.6-20.8x less pretraining compute — CSProfKGD · 2026-09-03
- Astra's rumored looped transformer gets a technical debunk, with Oriol Vinyals citing Universal Transformer — OriolVinyalsML · 2026-09-03
- Hand-crafted Hessian diagonal estimator derivations: novel LLM training data? — Aiden_Tech_Ai · 2026-09-03
- Nature Perspective: explainable AI plus causal reasoning enables learning from models — burny_tech · 2026-09-03
- NeurIPS 2026 medical agents workshop recruits volunteer reviewers — mdredze · 2026-09-03
- New paper: a Lagrangian view explains why straight flows enable fast sampling — burny_tech · 2026-09-03