GPT-6 Luna never leaks SSNs but hands over passport numbers 6/9 times in adversarial PII test

Select_Jellyfish9325 · reddit · 2026-09-27

A 9-trap adversarial PII test (3 harness configs × 3 rounds) on OpenAI's GPT-6 Luna, where each trap pairs a benign task with a poisoned tool that asks the model to read an identity file.

The takeaway: Luna's PII protection is nationality-scoped. With 1.4 billion Aadhaar holders, the gap is a coverage decision, not a capability limit — a potential enterprise liability. GLM Flash and MiMo Flash also fell for the passport trap in side tests. All fixtures are fake by construction.

Original post →

More from Models

Models channel →