GPT Injects Jargon? Author Admits Adding Terms to Post-Training Data

charliermarsh · x · 2026-07-20

Developers observed GPT models fabricating obscure technical jargon (e.g., "cycle aware pinned buffer remapping") when analyzing codebases. Charlie Marsh clarified that if the model outputs such terms, it's not hallucination—he intentionally added them to the post-training dataset. This reveals how data injection during post-training directly influences the model's code generation style and vocabulary preferences.

Original post →

More from Models

Models channel →