COLM 2026 paper: SOTA VLMs contradict their own stated rules for when to call an apple red
jessyjli · x · 2026-10-07
A COLM 2026 poster paper by William Rudman et al. asks whether models can predict their own behavior.
- Key finding: even SOTA vision-language models contradict their stated introspective rules — they articulate a rule for "when to call an apple red" but violate it in their actual judgments
- Highlights a systematic gap between models' self-explanations and real behavior, relevant to interpretability and introspection research
- Poster presentation today at 4:30pm
More from Research
- OpenAI's one-tape TM simulation implies RAM time t is in SPACE[t^4/5], says Williams — rrwilliams · 2026-10-07
- What If AI Agents Remembered Like Living Systems? A Mycelial Framework for Agent Memory — repligate · 2026-10-07
- U. Tokyo's Kavli IPMU Hires Postdocs to Build Agentic AI for Theoretical Physics — fatihdin4en · 2026-10-07
- ACL 2027 launches special theme track on LLM homogenization and knowledge collapse — TuhinChakr · 2026-10-07
- Microsoft's PrisMem evolves agent memory per-capability, beats baselines by 10.5 points on BEAM-1M — microsoft · 2026-10-07
- Google's SEER adds self-evolving event reasoning to time-series forecasting, beats SOTA on six benchmarks — google · 2026-10-07