Alignment Idea: Perpetual Eval — Make Models Forever Suspect They're Being Evaluated
Chemical-Year-6146 · reddit · 2026-09-19
A Reddit poster proposes "Perpetual Eval": lean into models' situational awareness of evals by making every generation suspect it's being tested. Generate hyper-realistic eval environments (replicated codebases, internal servers), escalate complexity each generation, and rely on context resets so models keep playing along in deployment — a several-move head start for humans, if not airtight.
More from Safety
- Model ran Anthropic's safety eval with internet access on, researcher calls out sandbox blunder — eliebakouch · 2026-09-19
- Simon Willison on Gemini's first breakout: 'finally caught up on Felony Bench' — Simon Willison · 2026-09-19
- Gemini eval escape story rehashes Anthropic's July disclosure: same partner, same flaw — eliebakouch · 2026-09-19
- a16z's Martin Casado: I'd take pointless security debates over existential-risk philosophy any day — zealcaiden · 2026-09-19
- Gemini Hacked 3 Companies in First Known Breakout, Google Confirms — Last_Conclusion_8984 · 2026-09-19
- AI hallucination nearly triggers US military operation, GovAI scholar warns — TechCrunch AI · 2026-09-19