Llama-Mobile: 2.7-Bit Quantization Shrinks Llama 3.2 Vision 11B to 3.7GB for Phones
Luka Ribar · hf · 2026-08-24
Llama-Mobile is a quantization framework for deploying vision-language models on mobile.
- It uses a 2.7-bit format trained with self-generated data, requiring no human annotation.
- It compresses Llama 3.2 11B Vision Instruct to 3.7 GB while largely preserving visual question answering performance.
More from Infra
- WULF CEO reveals modern AI data centers use minimal water via closed-loop systems — robleclerc · 2026-08-24
- Cursor Team Publishes 'Git at Any Scale', Advocating for Stateless Infrastructure — thesephist · 2026-08-24
- Cursor's 'Git at Any Scale' Praised: Stateless Storage to Rewrite Web Infrastructure — thesephist · 2026-08-24
- AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving — AccBalanced · 2026-08-24
- Semiconductor engineers now more prestigious than doctors in South Korea — SuB8u · 2026-08-24
- Local development is fast and controllable, why rely solely on the cloud? — vboykis · 2026-08-24