CLM-8B matches Jev on computer-use and tool-calling benchmarks while running up to 9x faster
Azaliamirh · x · 2026-09-24
A zero-shot evaluation thread comparing CLM-8B with Jev: across computer-use, gaming, and tool-calling tasks, CLM-8B matches Jev while running up to 9x faster. Speedups are largest with many candidates (WikiRacing) or when actions can be reused across states (T-Rex Game). CLM-35B, with better generalization and even greater speedups, ships early next month.
More from Models
- Why Jev may threaten frontier labs more than DeepSeek: an API so cheap everyone finds the waste — jobergum · 2026-09-24
- ChatGPT reportedly removes message cap on GPT-5.6 Luna for free users (unverified) — Aiden_Tech_Ai · 2026-09-24
- Grok 4.7 enters AutoResearchExam live leaderboard, ranks No.3 at 30min and No.4 after 24h auto-research — AlexGDimakis · 2026-09-24
- Researcher _xjdr: not liking astra, may go back to 5.6, eyeing Opus 5.5 and dsv4.1 flash — _xjdr · 2026-09-24
- OpenAI's MentalHealthBench scores clinicians below most AI models — and that reveals a flaw — r0ck3t23 · 2026-09-24
- Bug-finding ability grows exponentially costlier across models, Paweł Huryn benchmark shows — garrytan · 2026-09-24