Laguna tests 2.75 and 3.25 bpw quantization with NVFP4 experts and FP8 KV cache
QuixiAI · x · 2026-07-27
A post about Laguna’s 2.75 bpw and 3.25 bpw quantization experiments outlines a REAP-based setup with no pruning.
Key details:
- Keep 23–30% of experts in NVFP4
- Put the remaining experts into EXL3-style 2/3-bit formats
- Use FP8 for the KV cache
The author says the next steps are adding vision support, pushing 2.75 bpw down to 2 bits, and running benchmarks such as Terminal-bench-2.1 and GPQA Diamond.
More from Infra
- HyperNexus routes Claude, GPT, Gemini, and local models through one MCP server — HyperNexusLLC · 2026-07-27
- Voice AI’s real call cost is more than minutes: STT, TTS, SIP and retries — decant338 · 2026-07-27
- Micron and Meta paper says Spark can slow down 38× when shuffle spills to SSD — dr_alphalyrae · 2026-07-27
- Inference businesses may be charging 7× to 15× more than renting a GPU — JoshPurtell · 2026-07-27
- CUDA benchmark shows 64M-number sum runs 20× faster on GPU than CPU — ctjlewis · 2026-07-27
- AI buildout pushes data-center financing into more creative capital structures — GaryMarcus · 2026-07-27