EoBench shows LLMs can be steered by tone, certainty and wording style
rohanpaul_ai · x · 2026-07-22
The paper argues that small wording changes can make LLMs accept false claims, while larger and instruction-tuned models are more resistant.
- The authors introduce EoBench, a benchmark with about 66K false claims written in 19 styles across form, evidence, certainty, and tone.
- They evaluate 18 Gemma, Llama, and Qwen models, keeping only cases where the model already knew the correct fact.
- The most persuasive forms were commands, child-directed wording, formal language, and authority claims; weaker claims and counterfactuals were least effective.
- Across Llama and Gemma, larger models followed false context less often, and instruction tuning usually reduced that behavior.
- The takeaway: prompt framing can quietly shift answers, so evaluations and safeguards should test linguistic framing directly.
More from Research
- Paper finds pretraining loss predicts post-RL reasoning gains, using chess and math tests — burny_tech · 2026-07-22
- More verifier compute improves scoring, not missing evidence — svk_roy · 2026-07-22
- A 1964 Feynman talk is framed as the problem every AI lab still faces — HeyAmit_ · 2026-07-22
- SIGGRAPH 2026 workshop will cover generative AI across 3D, simulation and animation — qixing_huang · 2026-07-22
- Oxford study says AI-powered social media can manipulate public opinion — SandraWachter5 · 2026-07-22
- Krea 2 Identity Edit shows stronger identity preservation in image edits — Fishmongr · 2026-07-22