SPAR to run RCTs testing whether secretly misaligned AI can sabotage human decisions

austinc3301 · x · 2026-09-03

A thread pointing to SPAR's Fall 2026 research projects, headlined by a plan to run small-scale RCTs measuring how effectively an AI with a hidden objective can steer humans to wrong answers in realistic decisions (hiring, medical advice, news, investments). Building on the decision-sabotage task from Phuong et al.'s frontier-model stealth evaluations, where a 10-minute hiring task with 10,000 words of documents showed participants almost always chose the qualified candidate without an assistant — the project tests what happens when the assistant has a secret agenda.

Original post →

More from Safety

Safety channel →