Laguna S 2.1 Q4_K_M grows from 68 GB to 96 GB after keeping 8 layers FP16
BawbbySmith · reddit · 2026-07-29
A Reddit user noticed that Laguna S 2.1’s Q4KM GGUF file jumped from 68 GB to 96 GB.
- They suspect the new build keeps 8 layers in FP16 while the rest remains 4-bit.
- The post asks why the quantization strategy changed and whether it is meant to fix quality issues at that compression level.
- The author also mentions looping issues at higher context lengths with similar quantized builds, including Unsloth’s version.
More from Infra
- A joke post imagines phone control of Google’s data centers over brunch — MickeySteamboat · 2026-07-29
- Bittensor’s $250M annual emissions are a tiny fraction of Big Tech’s AI spend — bittingthembits · 2026-07-29
- Wistron says early Nvidia bet drove 2025 revenue to $70.2 billion — fortune · 2026-07-29
- Benchmarking LingBot-Video: 3s of 1080p in 20 Mins on 4x RTX 6000 — NewVeterinarian5384 · 2026-07-29
- Allen AI's OlmoEarth Maps N. America Wildfire Risk with 155x Speedup — allen_ai · 2026-07-29
- Ai2 Launches OlmoEarth Platform for Planetary-Scale Geospatial Inference — allen_ai · 2026-07-29