mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB
Spiritual_Impress_30 · reddit · 2026-09-30
A user thanks quantizer mradermacher: using their high-quality quants, gemma 4 26b hits 75 tok/s text generation and 1500 pp via LM Studio serving on 2x RTX 4060 8GB — "best IQ quants in the biz."
More from Infra
- South Korea expects record-high tax revenue amid AI chip boom — Polymarket · 2026-09-30
- Government weather model WRF ported to GPUs, running an order of magnitude faster with 250m fog forecasts — Scobleizer · 2026-09-30
- Ubuntu Snapdragon Edition ISO now available for download ahead of official 2027 release — carrycooldude · 2026-09-30
- Canonical and Qualcomm to bring official Ubuntu to Snapdragon X2 laptops in 2027, targeting local agentic AI — carrycooldude · 2026-09-30
- SortedRL: Microsoft Research tackles 70-74% GPU idle time in LLM reinforcement learning — burkov · 2026-09-30
- Intel's bizarre comeback: yields reportedly near 80%, raising talk of TSMC talent defections — 2C_ornot2C · 2026-09-30