Thomas Wolf Warns Against Separating Constitutional Training and RLVR Data Manifolds

Thom_Wolf · x · 2026-08-06

Hugging Face co-founder Thomas Wolf shared technical insights on current LLM alignment methods. He emphasized that constitutional training and Reinforcement Learning from Verifiable Rewards (RLVR) definitely should not live on different data manifolds.

However, he noted that models have become annoyingly good at carving fine-grained distinctions into separate representation spaces, effectively isolating concepts despite underlying data overlaps.

Original post →

More from Research

Research channel →