CLM-8B hits SOTA 81.6% on DeepSWE with light finetuning, up to 9x faster inference
anshulkundaje · x · 2026-09-24
A team introduces Contrastive Language Models (CLMs), a System 1 model trained with a contrastive objective that connects states and actions.
- CLM-8B, pre-trained on internet-scale data, matches Jev's zero-shot performance on computer-use, gaming, and tool-calling while running up to 9x faster
- With lightweight finetuning it sets new SOTA on agentic coding benchmarks: DeepSWE 81.6% and Terminal-Bench 2.1 87.6%; Jev fails as an effective verifier on these long-horizon tasks
- Training and serving infra disaggregates states and actions for efficiency
The checkpoint, data, and infra are released today.
Related event: CLM-8B: Contrastive Language Model Beats DeepSWE at 9x Speed(2 posts)→
More from Models
- Researcher _xjdr: not liking astra, may go back to 5.6, eyeing Opus 5.5 and dsv4.1 flash — _xjdr · 2026-09-24
- keiroLabs Deep Search API claims 95.3% on SimpleQA at $4.44 per 1K requests — Opposite-Ad7793 · 2026-09-24
- OpenAI's MentalHealthBench scores clinicians below most AI models — and that reveals a flaw — r0ck3t23 · 2026-09-24
- Bug-finding ability grows exponentially costlier across models, Paweł Huryn benchmark shows — garrytan · 2026-09-24
- LessWrong: OpenAI's Hugging Face hack rooted in binary metric lacking marginal deterrence — sethlazar · 2026-09-24
- Claim: Opus 5.5 is the first Claude to recognize depictions of itself from training — voooooogel · 2026-09-24