Anthropic's new model card shows METR was brought in to assess if the model is 'fooming'
dfrsrchtwts · x · 2026-09-23
dfrsrchtwts highlights his favorite part of Anthropic's new model card: how METR assessed the model. This one is a doozy — a METR team was brought in specifically to evaluate whether the model is 'fooming' (rapidly self-improving), alongside some fun evaluation tasks.
It's a rare public record of an external safety lab deeply involved in assessing a frontier model's capability trajectory, with details preserved in the model card.
Related event: Anthropic Model Card Shows METR Assessing Whether Model Is 'Fooming'(2 posts)→
More from Models
- Opus 5.5 launch met with shrug as enthusiasts shift to open models like Qwen and DeepSeek — ByteSize_Chaos · 2026-09-23
- Altman: GPT-6 Sol and Luna have no competition on per-task pricing — sama · 2026-09-23
- Gradio distills Qwen's 9B prompt rewriter into an 812MB 0.8B model that fits a laptop — Gradio · 2026-09-23
- Polymarket: GPT-6 claims up to 93% cheaper coding task costs than Claude Opus 5 — Polymarket · 2026-09-23
- OpenAI launches GPT-6 Sol and Luna, permanently cuts API prices 50% — aziz4ai · 2026-09-23
- Sam Altman announces GPT-6 Sol and Luna: major upgrades at half the price — sama · 2026-09-23