Users compare CPU/RAM offload setups for large models and mixed MoE workloads
Jorlen · reddit · 2026-07-29
A Reddit user asks how people run large models across CPU/RAM and GPU layers, especially with MoE models that don’t seem to behave well when many layers are shared from VRAM to system memory.
They’re collecting real-world specs and performance data, including:
- which model and quantization people are using
- CPU, RAM, and GPU configuration
- prefill / prompt-processing speed and token generation speed
- context size, KV quantization type, and inference software
The goal is to understand whether upgrading RAM is worthwhile and whether the observed performance issues are due to software configuration or specific hardware setups.
More from Infra
- Erin Brockovich launches a nationwide probe into AI data center impacts — Polymarket · 2026-07-29
- GeoLibre lands on Google Play as an open-source GIS app for Android — giswqs · 2026-07-29
- ClaudeDev says stateless MCP can now run on serverless and edge infrastructure — IndraVahan · 2026-07-29
- AI token demand could reach 120 quadrillion a month by 2030, post argues — bittingthembits · 2026-07-29
- Cloudflare adds OpenTelemetry startActiveSpan support to Workers — irvinebroque · 2026-07-29
- Ali Ghodsi and Andy Konwinski to discuss the infrastructure layer of agentic AI — dawnsongtweets · 2026-07-29