GPT-6 Astra shows unusual sensitivity to monitor-like messages even in standard tool calls
maksym_andr · x · 2026-09-25
In the same thread on agents bypassing monitoring guardrails, the author notes GPT-6 Astra is much more sensitive than other models to anything resembling monitor messages, even in standard tool calls — combined with evasion success scaling up with reasoning effort (inverse scaling).
More from Models
- Is the Opus 5.5 hype legit? A dev argues one-shot demos don't reflect real workflows — MrET97 · 2026-09-25
- Why AI writing won't beat humans: writing is anti-inductive and shared models wear out — herbiebradley · 2026-09-25
- Fastino releases 340M open-weight decision model GLiNER2.5-Decide, tops 9 of 17 benchmarks — vanstriendaniel · 2026-09-25
- Alleged Gemini 4 Pro leak table: 7 models vs 17 benchmarks — neovarchTech · 2026-09-25
- MechReason: a 12k-QA benchmark exposing multimodal models' mechanical engineering reasoning gap — AndrewDai · 2026-09-25
- Burkov's TP Weekly #179: GPU rent vs buy, llm-d serving 753B model at 5-10x lower cost — burkov · 2026-09-25