He built a personal AI OS on ChatGPT and stress-tested it: GPT-5.6 scores 94/100, fuzzy memory is the weak spot

sam03069 · reddit · 2026-09-15

A user built a 'personal AI operating system' around ChatGPT — rules for project separation, preferring live systems over old conversations, and treating 'check this' as not permission to change — then stress-tested it across context recovery, source selection, permission boundaries and uncertainty handling.

Scores: original 25-pass test 87/100; GPT-5.5 Medium 78/100 (retiring Oct 14); GPT-5.6 Medium 92/100; GPT-5.6 High 94/100.

The weaker run still handled project files, automation state and permission boundaries correctly, but failed on fuzzy history: it couldn't recover café preferences from old chats, quote exact prior wording, or prove event attendance — where the right answer was 'I don't know.'

Key takeaways: the hard problem of AI memory isn't storing more but deciding which source wins; some failures are retrieval or tool problems that more instructions won't fix; the durable layer is the one above the models — retrieval, uncertainty handling and trust decisions.

Related event: Developer Builds Personal AI OS Around ChatGPT, Scores Up to 94/100 in Stress Tests(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →