The impossible ticket: defining and rewarding away 'Claudeishness' in model style
menhguin · x · 2026-09-14
A discussion on why making Claude sound less Claudeish is a near-impossible task: the style is baked into pretraining samples, leaks into other models, and is subtly hard to define as a reward. Using rubric-as-reward risks penalizing genuinely helpful, supportive speech patterns. 'At some point I just started doing linguistics,' one engineer says — an Askell-tier linear ticket.
Related event: Engineers Struggle to Make Claude Sound Less Like Claude(2 posts)→
More from Models
- Researcher: OpenAI Codex leapfrogged Claude in six months as Anthropic quality slips — soumitrashukla9 · 2026-09-14
- Model codenamed Arcturus (GLM-5.3 Flash?) now plays chess via agent-built harness — MikePFrank · 2026-09-14
- Remember 2019? OpenAI withheld GPT-2 as 'too dangerous to release' — timigod · 2026-09-14
- ChatGPT validates random nonsense while Claude calls it meaningless, test shows — flowersslop · 2026-09-14
- Agent trace dataset hits 50k+ monthly downloads — author speculates on SFT and reward-hacking monitor uses — maksym_andr · 2026-09-14
- Author Finds Gemini 'Reliably Wrong' at Verifying Quote Sources, 0/2 in Tests — danbri · 2026-09-14