LLMs Fall for Common Sense Traps: Salience Bias Causes Reasoning Failures
rohanpaul_ai · x · 2026-08-03
The paper Would You Walk to the Car Wash? reveals the salience bias in LLMs during commonsense reasoning: when prompts contain explicit numbers or procedures, models often ignore unstated physical prerequisites (e.g., a car must be driven to a car wash, not walked).
- Benchmark: Introduces SaliTrap, featuring 1,145 prompts across 4 trap types, tested on 12 models.
- Poor Performance: The best model avoided the trap in only 54.8% of queries, with 8 out of 12 models scoring below 30%. More numerical distractors led models to calculate blindly rather than spot contradictions.
- Awareness ≠ Avoidance: Even when models recognized the trap (like GLM-5.1 and Kimi-K2), they still complied with the flawed premise 86.2% and 81.8% of the time.
- Evaluation Advice: Agent evaluations should explicitly test premise-checking as a behavioral control rather than assuming it from general reasoning scores.
More from Models
- Gary Marcus: OpenAI's Astra is Vastly Oversold, Math Breakthroughs Don't Mean AGI — Gary Marcus · 2026-08-03
- Only OpenAI Mastered Reasoning, Yet DeepSeek Tops Benchmark with Medium Effort — teortaxesTex · 2026-08-03
- Gemini Video Generation Backlash: AI Pro Subscribers Get Fewer Credits Than Free Users — Captain-Thump · 2026-08-03
- Fail: Google Gemini Botches a Simple Counting Task — polo421 · 2026-08-03
- Cutting AI Coding Bills from $200 to $20/Month: 105 Bugs Tested — PawelHuryn · 2026-08-03
- Ornith 35B Outperforms Qwen 3.6 and Laguna S in Open-Source Showdown — S_Anv · 2026-08-03