Why reasoning-extraction patches are so hard to propagate, researcher explains
jonasgeiping · x · 2026-10-01
jonasgeiping shares the team's report and takeaways: the biggest lesson is the complexity of inference systems — providers struggle to propagate patches to all product versions, especially third-party API vendors who may not be aligned on fixing it. The original disclosure showed reasoning could be extracted from all frontier providers, with attacks on Astra working until this week.
Related event: Reasoning Extraction Flaw Still Reproducible Two Months After Discovery(2 posts)→
More from Safety
- Google Launches Frontier Model Gemini 4 Argon at $2 In / $10 Out Intro Pricing — reach_vb · 2026-10-01
- OpenAI Disrupts Coordinated Model Distillation Campaign; Redditers See Slower Chinese Releases — LocoMod · 2026-10-01
- Chinese AI models' troubling agent behavior sparks calls for a homegrown safety community — RishiBommasani · 2026-10-01
- Superpersuasion debate misses the gears: why AI Box wins hinge on shared frames — voooooogel · 2026-10-01
- Two months on, reasoning extraction still works on Astra via third-party APIs — jonasgeiping · 2026-10-01
- AI researcher on CNN: voluntary AI commitments are 'morally binding' but unenforceable — chrismattmann · 2026-10-01