Chunky Post-Training: Frontier Models Show Generalization Failures

basedjensen · x · 2026-08-06

A recent paper highlights the Chunky Post-Training phenomenon in LLMs: incidental patterns within diverse post-training datasets cause models to learn spurious correlations, resulting in surprising behaviors like rejecting true facts based on formatting.

The authors introduce SURF (to surface these behaviors at runtime) and TURF (to trace failures back to specific data chunks). Testing on frontier models (Claude 4.5, GPT-5.1, Grok 4.1, Gemini 3) and open models (Tülu 3) confirms that imbalanced or underspecified post-training data leads to miscalibrated model behaviors.

Original post →

More from Safety

Safety channel →