Train Only Projector to Add New Modalities Without Regressing LLM Capabilities
rohanpaul_ai · x · 2026-08-25
The paper 'Projector Is All You Train' demonstrates that adding a new modality (e.g., 3D) to an LLM requires training only the projector mapping encoder outputs to the embedding space, leaving the backbone frozen. This approach trains twice as fast as joint fine-tuning and matches or outperforms it on 3D tasks. Conversely, joint fine-tuning caused the Llama backbone's GSM8K accuracy to collapse from 86.96% to 0.61%.
More from Models
- GLM-5.3 shows strong cost-efficiency vs GPT 5.6 on TerminalBench-3.0 — zainhas · 2026-08-25
- Grok Voice tops Speech-to-Speech Index, outperforming all GPT Realtime models — XFreeze · 2026-08-25
- Redditor finds MiniMax H3 performs far better with Mandarin prompts than English — apoke890 · 2026-08-25
- Suspected Gemini 3.8 Flash leak surfaces on Reddit — Last_Conclusion_8984 · 2026-08-25
- User swaps one word, gets very different AI results — claims sexism — RecordingThis6802 · 2026-08-25
- Fal releases post-trained H3 model co-optimized with custom inference stack — isidentical · 2026-08-25