LongCat-Flash-Lite-Sparse and Qwen Uncensored Models Released in GGUF

LLMFan46 · reddit · 2026-08-30

Community modeler LLMFan46 released several quantized models, headlined by LongCat-Flash-Lite-Sparse. This 69B-A3B model adds sparse attention and 1M context length support over the original, requiring a custom llama.cpp fork. The release also includes Ultra Uncensored Heretic versions of Qwen3.8-27B and Qwen3.5-122B-A10B, plus Qwen3-Coder-Next and Laguna-S2.1 with vision. Formats include GGUF, Safetensors, GPTQ-Int4, and NVFP4.

Original post →

More from Infra

Infra channel →