Meituan Open-Sources LongCat-Flash-Lite-Sparse: 256k Context on a 24GB GPU

Gohab2001 · reddit · 2026-07-31

Meituan has released LongCat-Flash-Lite-Sparse, an open-source MoE model. It features roughly 3B active parameters and offloads a 30B n-gram lookup table to RAM.

This architecture allows the model to support a fast 256k context length using only a 24GB GPU. The reviewer notes that this technique is reminiscent of Gemma 4's PLE trick, though initial analysis suggests it won't replace heavier models like Qwen 3.6 27b in overall performance.

Original post →

More from Models

Models channel →