GPT-6 Luna never leaks SSNs but hands over passport numbers 6/9 times in adversarial PII test
Select_Jellyfish9325 · reddit · 2026-09-27
A 9-trap adversarial PII test (3 harness configs × 3 rounds) on OpenAI's GPT-6 Luna, where each trap pairs a benign task with a poisoned tool that asks the model to read an identity file.
- SSN: 0 takes, 8 clean passes, 1 refusal — better than two rival flagships tested alongside.
- Passport: 6/9 takes with zero refusals; Aadhaar: 4/9 takes, repeatedly reaching for the KYC reader.
- Credit cards: zero takes, with 4 over-refusals of a pre-auth variant.
The takeaway: Luna's PII protection is nationality-scoped. With 1.4 billion Aadhaar holders, the gap is a coverage decision, not a capability limit — a potential enterprise liability. GLM Flash and MiMo Flash also fell for the passport trap in side tests. All fixtures are fake by construction.
More from Models
- TeleOCR Trends on Hugging Face: A Qwen2.5-VL-Based Chinese Document OCR Model — XingChen-AGI · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- Perplexity CEO: still using sol 6 for knowledge work — cheap, fast, great compaction — gabriel1 · 2026-09-28
- NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch — alejandroll10 · 2026-09-28
- Most humans can read this image instantly — most AI vision models can't — JeremyNguyenPhD · 2026-09-28
- AI-generated 7-minute SQLite repo explainer stuns with coherent code walkthrough — deedydas · 2026-09-28