Extra 30% Training FLOPs Yields Inference Performance of 2x Compute

teortaxesTex · x · 2026-08-14

The post quotes developer @xidulu on a novel model training paradigm. By incorporating (soft, free) latent feedback decoding and fused double prefill + soft decoding, this method requires only an extra 30% of FLOPs during training. However, it enables the model's free-form generation performance to match that of models trained with double the compute using traditional methods.

Original post →

More from Research

Research channel →