Respan launches Span-1, a 4B eval model claimed to beat frontier LLM-as-a-judge
HeyAmit_ · x · 2026-09-25
RespanAI introduced Span-1, a purpose-built 4B model for evaluation, claiming it beats frontier models at LLM-as-a-judge. The pitch: a judge must understand the entire trace — messages, tool calls, evidence, metadata — and traces contain instructions you shouldn't trust. If a 4B model does this better, the author argues, current eval practices may be backwards.
More from coding & agent
- ECHO Paper Lands NeurIPS Spotlight: Terminal Agents Learn World Models for Free, Doubling GRPO Pass@1 — DimitrisPapail · 2026-09-25
- Phoson: A Framework-Free Minimal Open-Source Agent Runtime With Terminal CLI — abelsr_1710 · 2026-09-25
- Kaigen, a C-based AI-native game engine, opens closed beta — gdechichi · 2026-09-25
- One-prompt Minecraft: AI-generated voxel game open-sourced, runs in browser and on Windows — gdechichi · 2026-09-25
- He used $100 of Claude credits to land his first VSCode PR — ThePeterMick · 2026-09-25
- What the OpenAI-Hugging Face incident says about agent oversight — rainerhahnekamp · 2026-09-25