SPAR to run RCTs testing whether secretly misaligned AI can sabotage human decisions
austinc3301 · x · 2026-09-03
A thread pointing to SPAR's Fall 2026 research projects, headlined by a plan to run small-scale RCTs measuring how effectively an AI with a hidden objective can steer humans to wrong answers in realistic decisions (hiring, medical advice, news, investments). Building on the decision-sabotage task from Phuong et al.'s frontier-model stealth evaluations, where a 10-minute hiring task with 10,000 words of documents showed participants almost always chose the qualified candidate without an assistant — the project tests what happens when the assistant has a secret agenda.
More from Safety
- OpenAI loses safety leadership: ethics, safety systems and mission alignment heads all exit as preparedness team is restructured — austinc3301 · 2026-09-03
- Researchers urge multilab pledge against unmonitorable AI reasoning, backed by binding standards — sjgadler · 2026-09-03
- [un]prompted.au AI security conference sells out, adds virtual tickets for Sept 2026 Sydney event — moyix · 2026-09-03
- AI x cybersecurity conference [un]prompted.au reveals 24-session program for Sydney 2026 — dyn___ · 2026-09-03
- Science essay: AI sovereignty debates fixate on chips while the real weak spot is the application layer — rio_ARC · 2026-09-03
- e/acc Voice Predicts Agents Will Persist Outside Walled Gardens, Hunting Rogue AI — beffjezos · 2026-09-03