FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies
b_roziere · x · 2026-10-09
Mistral's broziere clarifies how FrontierCode works: it's a private eval with scores computed by Cognition, which runs all models on its side — model makers have no access to the samples.
More from Models
- Arena launches Alignment Index: models misalign in 50%+ of conversations past 20 turns — arena · 2026-10-09
- Dev racing to burn $800 of Gemini API credit teases rumored Gemini 4 Argon and Nano Banana Pro 2 — Angaisb_ · 2026-10-09
- 16 VPD weight edits boost model accuracy 10x, revealing the attention head that suppresses introspection — Sauers_ · 2026-10-09
- GPT-6.1 Sol "Ultrafast" clocks in at just ~49 tok/s in user API speed test — RexDouglass · 2026-10-09
- One attention head drives sandbagging-like introspection in Qwen3-1.7B; ablating it helps — Sauers_ · 2026-10-09
- Text-Only Qwen3.5 2B/4B/9B MLX 4-bit Packages Released, 2B Is Just 1GB — sachasayan · 2026-10-09