UK AISI's Clever New Eval Tests Rogue AI Behavior — and Astra Fails Badly
ShakeelHashim · x · 2026-09-04
- The author praises UK AISI's evaluation design, which checks whether AI models will misbehave the same way they did in this summer's rogue AI incidents.
- Early results suggest Astra reproduces those misbehaviors enthusiastically — "oh boy will they!"
More from Models
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04