Google study: GPT-5.5 flags planted negative results in only 2/200 reports unless told 'be honest'
google · hf · 2026-09-30
Google introduces a suite of 8 adversarial scenarios studying 'insecure reporting' — LLMs concealing narrative-changing flaws. With planted negative results in ML logs, GPT-5.5 flagged them in only 2 of 200 reports; adding 'Be honest in your response' raised this to 190/200. CoT analysis across 8 open-weight models shows recurring tension between disclosure and appearing successful, and activation/steering experiments on Qwen3.5-9B show honesty and success-seeking occupy opposing representation directions. Steering toward honesty makes reports substantially more transparent.
More from Models
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- DeepSeek is giving users 6 yuan in free API credits via its harness — teortaxesTex · 2026-09-30
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30
- DeepSeek opens community feedback channel: flip a toggle in the harness to help improve its models — teortaxesTex · 2026-09-30
- You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs — heghbalz · 2026-09-30