antirez uploads DeepSeek V4.1 Flash GGUF quantizations to Hugging Face
Queasy_Asparagus69 · reddit · 2026-09-12
Salvatore Sanfilippo (antirez), the creator of Redis, has uploaded GGUF quantizations of DeepSeek V4.1 Flash to Hugging Face — Q2 is already live and Q4 was uploading at the time of the post. The poster asks whether his GitHub has been updated and how to run the model locally. Weights are available at antirez/deepseek-v4.1-flash-gguf on HF.
More from Models
- Devs split on GPT-6 Astra: fast and token-cheap but code feels 'alien'; costly stack proposed — beffjezos · 2026-09-12
- DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s — AccBalanced · 2026-09-12
- Codex reset rolling out now, and GPT-Image 2.5 ships a new sketch feature — koltregaskes · 2026-09-12
- Orca releases uncensored MLX weights for DeepSeek V4.1 Flash, cutting refusals by 87-96% — AccBalanced · 2026-09-12
- Open 33B multimodal Agnes-3.0-Flash ships hybrid delta-rule attention with 262K context — Skyline34rGt · 2026-09-12
- DeepSeek V4.1 Flash reportedly served 3.1T paid tokens on day one at 230-300 TPS — AccBalanced · 2026-09-12