GPT Injects Jargon? Author Admits Adding Terms to Post-Training Data
charliermarsh · x · 2026-07-20
Developers observed GPT models fabricating obscure technical jargon (e.g., "cycle aware pinned buffer remapping") when analyzing codebases. Charlie Marsh clarified that if the model outputs such terms, it's not hallucination—he intentionally added them to the post-training dataset. This reveals how data injection during post-training directly influences the model's code generation style and vocabulary preferences.
More from Models
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11