Model Behavior Hinges on Post-Training

gerardsans · x · 2026-07-19

The discussion emphasizes that behavioral differences in models stem not just from pre-training data and architecture, but also from post-training and RLHF. The author points out that while many labs share common training corpora, the exact composition is opaque; furthermore, the data and rules used during post-training are often kept strictly behind closed doors. Therefore, evaluating model performance requires looking beyond pretraining to include post-training and the alignment process.

Related event: Claude's Constitutional AI: Post-Training Dictates Model Behavior(2 posts)→

Original post →

More from Research

Research channel →