GPT Injects Jargon? Author Admits Adding Terms to Post-Training Data
charliermarsh · x · 2026-07-20
Developers observed GPT models fabricating obscure technical jargon (e.g., "cycle aware pinned buffer remapping") when analyzing codebases. Charlie Marsh clarified that if the model outputs such terms, it's not hallucination—he intentionally added them to the post-training dataset. This reveals how data injection during post-training directly influences the model's code generation style and vocabulary preferences.
More from Models
- Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks — ChrisGPT · 2026-07-22
- Post says Google DeepMind has gone over a year without pretraining a new base model — teortaxesTex · 2026-07-22
- Current setup is 8,192 input tokens and 2,048 output tokens, with 8k/512 next — TheZachMueller · 2026-07-22
- Kimi K3 feels slower than K2.7, but stronger on long coding jobs and refactoring — Far-Presence2711 · 2026-07-22
- Poolside launches Laguna S 2.1 with 118B parameters and 8B active per token — Madisonkanna · 2026-07-22
- What are the best models to run on 48 GB of VRAM with two RTX 3090s? — ludos1978 · 2026-07-22