He built a personal AI OS on ChatGPT and stress-tested it: GPT-5.6 scores 94/100, fuzzy memory is the weak spot
sam03069 · reddit · 2026-09-15
A user built a 'personal AI operating system' around ChatGPT — rules for project separation, preferring live systems over old conversations, and treating 'check this' as not permission to change — then stress-tested it across context recovery, source selection, permission boundaries and uncertainty handling.
Scores: original 25-pass test 87/100; GPT-5.5 Medium 78/100 (retiring Oct 14); GPT-5.6 Medium 92/100; GPT-5.6 High 94/100.
The weaker run still handled project files, automation state and permission boundaries correctly, but failed on fuzzy history: it couldn't recover café preferences from old chats, quote exact prior wording, or prove event attendance — where the right answer was 'I don't know.'
Key takeaways: the hard problem of AI memory isn't storing more but deciding which source wins; some failures are retrieval or tool problems that more instructions won't fix; the durable layer is the one above the models — retrieval, uncertainty handling and trust decisions.
More from AGI Musings
- OpenAI capabilities researcher Dan Selsam makes public statement on AI risk — connoraxiotes · 2026-09-15
- Kai-Fu Lee launches 'AI Native' book on enterprise AI transformation — kaifulee · 2026-09-15
- Philosopher pushes back: denying AI rights implies rejecting computational theory of mind — dioscuri · 2026-09-15
- Bengio took years to be convinced AI risk concerns were worth taking seriously — S_OhEigeartaigh · 2026-09-15
- Hype vs. real: why both camps reading AI-lab doomsday statements are right — TotalPhilanthrope · 2026-09-15
- mark_k's Quip: What Scares Me Most Is a Future Controlled by Doomers — mark_k · 2026-09-15