LongCat-Flash-Lite-Sparse and Qwen Uncensored Models Released in GGUF
LLMFan46 · reddit · 2026-08-30
Community modeler LLMFan46 released several quantized models, headlined by LongCat-Flash-Lite-Sparse. This 69B-A3B model adds sparse attention and 1M context length support over the original, requiring a custom llama.cpp fork. The release also includes Ultra Uncensored Heretic versions of Qwen3.8-27B and Qwen3.5-122B-A10B, plus Qwen3-Coder-Next and Laguna-S2.1 with vision. Formats include GGUF, Safetensors, GPTQ-Int4, and NVFP4.
More from Infra
- DLSS 5 Video Player tested: Runs slow on 3090 Ti due to FP8 lack — fallengt · 2026-08-30
- Data Centers Drive US Reindustrialization: Benefits from Taxes to Jobs — GavinSBaker · 2026-08-30
- Running Qwen3.8-Next-Flash on 96GB RAM: offload the n-gram table to SSD — Iory1998 · 2026-08-30
- Vietnam's OneNexus quantizes GLM-5.3 to MXFP4 — xiaosun86 · 2026-08-30
- Pi Agent setup: Using Qwen 27B as an Oracle for co-development — Thrumpwart · 2026-08-30
- Full-stack AI EDA to disrupt chip design economics — ai · 2026-08-30