Can older MXM GPUs run lightweight LLMs?
Big_black_click · reddit · 2026-07-21
The author asks what support exists for running lightweight LLMs on older MXM laptop GPUs such as the GTX 980M or Quadro P5000.
They want to buy an 8GB MXM card for home automation, but note that many newer MXM cards use proprietary DGFF connectors and lack adapters or pinout information, making them hard to use.
The post also asks:
- which frameworks or engines still support older laptop GPUs with older CUDA versions
- whether vLLM is usable in this scenario
- how to convert a Llama 3.2 3B model to GGUF, noting issues with the conversion script and Meta model card hashes
Overall, it is a practical discussion about squeezing small models onto legacy consumer GPU hardware.
More from Infra
- NVIDIA publishes Vera CPU architecture details before AMD’s AI event — ryanshrout · 2026-07-22
- oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads — awnihannun · 2026-07-22
- Strangeworks launches Aura to turn enterprise ops into production optimization systems — whurley · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- HilbertRaum open-sources a fully local AI chat and document analysis app for private use — Vladowski · 2026-07-22
- Hybrid and local inference are emerging as a response to AI energy and token costs — dmitry140 · 2026-07-22