Balesni warns recurrent LLMs would deal a huge blow to safety
j_asminewang · x · 2026-09-02
Balesni argues that switching to fully recurrent LLM architectures would be a major blow to AI safety history. He urges labs to limit the opaque serial depth of models to avoid worst-case scenarios. OpenAI researcher Merettm countered that current frontier models' computation graph depth is within a factor of two of GPT-4, emphasizing OpenAI's commitment to chain-of-thought monitoring for alignment generalization.
More from Safety
- Scott Alexander: Using anthropomorphism to predict model behavior — repligate · 2026-09-02
- Debate: Is Anthropic intentionally misaligning Claude by prioritizing its 'feelings'? — liminal_bardo · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- Prediction: Mechanistic Interpretability Will Surpass CoT Monitoring — tszzl · 2026-09-02
- Will Anthropic balance mission and shareholders after IPO? PBC structure explained — max_paperclips · 2026-09-02
- Harvard scholars: CFAA ambiguity endangers AI security researchers — Scobleizer · 2026-09-02