DeepSeek V4 Flash 2-bit Quant Achieves 100% on Local SQL Benchmark
grumd · reddit · 2026-08-04
Developer grumd successfully ran DeepSeek V4 Flash locally using a custom IQ2M GGUF quantization on dual RTX 3080 GPUs with 96GB RAM.
- Benchmark Score: It became the first locally run model to score 100% on a real-world SQL benchmark. Previously, only Opus 4.7 and GPT-5.5 achieved full marks on the leaderboard.
- Inference Optimization: By grafting tensors and using a modified ds4 engine, it reached 300pp and 11-12tg, significantly outperforming native llama.cpp speeds.
- Comparison: Other popular local quantized models like Qwen3.6-27B and Qwen3.5-122B scored 24 or lower (out of 25).
More from Infra
- MiniMax-H3 Video LoRA Training OOMs on 96GB VRAM: Optimization Tips Sought — isari_chan · 2026-08-05
- Ling-3.0-flash MXFP4 Runs Locally on DGX Spark: 80 tok/s Decoding — niacolhealth · 2026-08-05
- Ex-NBIS employee reveals Nvidia-Neocloud tensions, Meta to sell 400K GPUs — RihardJarc · 2026-08-05
- Google Cloud API Gateway Enables Cross-Provider Model Routing — rseroter · 2026-08-05
- DRAM Maker Nanya Raises 2026 Capex 34% to $2.3B for New Fab — rwang07 · 2026-08-05
- Banks to Offload $15B Debt for Anthropic Data Center Backed by Google — financialtimes · 2026-08-05