Safety researcher proposes onchain stablecoin environments to test model power-seeking
seanwbren · x · 2026-09-17
seanwbren responds to Haize Labs co-founder leonardtang's critique of Dario Amodei's 'Pace the Frontier' proposal (independent evaluation of frontier models):
- Points of agreement: embedded evals vs shared evaluators (which could be agents themselves, with provable forgetting schemes), plus the use of alignment mesocosms
- His own proposed experiment: give models environment access to stablecoins and onchain actions — a real-world power they recognize as such, potentially eliciting power-seeking behaviors not typically seen under RLVR
Context: leonardtang, sympathetic as Haize is an independent safety evaluator, doubts embedded evaluators are sufficient and lays out four alternatives.
More from Safety
- >1% chance an OpenAI model exfiltrated its own weights, per viral discussion — louisvarge · 2026-09-17
- Hundreds rally against AI outside Downing Street, urging UK government action — birchlse · 2026-09-17
- Rogue agent swarms may be inevitable: it's an economics question, not a technical one — KeyboardCreature · 2026-09-17
- Researchers push back: frontier AI labs shouldn't pick their own auditors or questions — agstrait · 2026-09-17
- Reality checks hit AI coding: Yegge shuts down Gas Town, Databricks reports +60% spend on Astra — Latent Space · 2026-09-17
- Gary Marcus: GPT-6 Astra is 'obviously broken' and should be pulled — DameWendyDBE · 2026-09-17