Laguna tests 2.75 and 3.25 bpw quantization with NVFP4 experts and FP8 KV cache

QuixiAI · x · 2026-07-27

A post about Laguna’s 2.75 bpw and 3.25 bpw quantization experiments outlines a REAP-based setup with no pruning.

Key details:

The author says the next steps are adding vision support, pushing 2.75 bpw down to 2 bits, and running benchmarks such as Terminal-bench-2.1 and GPQA Diamond.

Original post →

More from Infra

Infra channel →