ECCV 2026 paper says some multimodal LLMs do worse with one-shot examples
HildeKuehne · x · 2026-07-29
Some multimodal LLMs get worse with one-shot examples, and the paper says class names are part of the reason
The post argues that on standard few-shot benchmarks, several MLLMs do better 0-shot than 1-shot. When semantic class names are removed, in-context prompting can nearly collapse.
The linked ECCV 2026 paper proposes DeCoDe and claims it explains why examples can hurt instead of help in some multimodal settings.
More from Research
- Real-to-Sim Work Explodes as New Tool Automates Robot Scene Import — chris_j_paxton · 2026-07-29
- LLMs on robots jump real-world success from 16.7% to 97.3% — tri_dao · 2026-07-29
- Simple teleoperation recordings plus soft compliance can solve more robot tasks than expected — ihorbeaver · 2026-07-29
- Studies find AI therapy replies often score higher on empathy than human clinicians — sapinker · 2026-07-29
- TransluceAI proposes oversight foundation models to catch reward hacking at scale — JacobSteinhardt · 2026-07-29
- Why tokenizers still resist end-to-end optimization despite years of pretraining — seanmcdonaldxyz · 2026-07-29