CoLT teaches multimodal models to think in latent steps and cuts inference 10.1x
jiqizhixin · x · 2026-07-23
Researchers from NTU and collaborators introduce CoLT, a method that makes multimodal models reason through a chain of latent thoughts instead of verbose text.
What it does
- Uses hidden representations as intermediate reasoning steps.
- Adds a lightweight decoder to supervise each latent step for stable training.
- Avoids the need for costly visual annotations used by some prior latent-reasoning approaches.
Reported gains
- Beats prior latent reasoning methods such as CODI and SIM-CoT.
- Cuts inference time by 10.1×.
- Reduces text decoding by 22.6×.
The post links the paper, code, and a report, positioning CoLT as a more efficient way for VLMs to think without generating long textual chains.
More from Multimodal
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11