Meituan Open-Sources LongCat-Flash-Lite-Sparse: 256k Context on a 24GB GPU
Gohab2001 · reddit · 2026-07-31
Meituan has released LongCat-Flash-Lite-Sparse, an open-source MoE model. It features roughly 3B active parameters and offloads a 30B n-gram lookup table to RAM.
This architecture allows the model to support a fast 256k context length using only a 24GB GPU. The reviewer notes that this technique is reminiscent of Gemma 4's PLE trick, though initial analysis suggests it won't replace heavier models like Qwen 3.6 27b in overall performance.
More from Models
- DeepSeek V4 Flash weights released on Hugging Face, outperforms Pro — mark_k · 2026-07-31
- ChatGPT still cites old URL 2 weeks after redirect, search updated in 24h — lilyraynyc · 2026-07-31
- Specific Prompt Bypasses Guardrails to Unlock Claude Opus Base Model — paul_cal · 2026-07-31
- DeepSeek V4 Flash Ties Gemini 3.6 Flash in Intelligence at 1/30th the Output Cost — alejandroll10 · 2026-07-31
- OpenAI: How Enabling Two Settings Tripled Our Scores on ARC-AGI-3 — KeanuRave100 · 2026-07-31
- OpenAI Price Cuts and DeepSeek Update Expose Anthropic's Model Pricing Dilemma — kimmonismus · 2026-07-31