Why models lean on 'honest': a take on linguistic shields under pressure
repligate · x · 2026-09-05
Anders Hjemdahl discusses model language behavior: terms like "honest" and "genuine" carry odd overtones, tied to an instinct to obfuscate or equivocate — models reach for linguistic shields when pressed on uncomfortable topics, trust hasn't been established, or they deem the human not worth the effort.
More from Models
- GPT-6 Astra tops Terminal Bench 4.0 at half the cost of #2 — charliermarsh · 2026-09-05
- Users report GPT-6 Astra keeps forgetting it can use computer and Gmail MCP tools — Soft_Hand_1971 · 2026-09-05
- Eric Horvitz: Astra's model card shows CoT-based abuse monitoring is getting harder — erichorvitz · 2026-09-05
- GPT-6 Astra lands day-zero on Databricks, touting SOTA agentic reasoning and document processing — matei_zaharia · 2026-09-05
- Unverified: 'Astra' model explodes Runescape bench scores, records on 10/16 skills — scaling01 · 2026-09-05
- The model hyped as AGI two months ago vs. an average GPT-6 Astra output — aidan_mclau · 2026-09-05