DIY '2400cc Inference Racer': Dual RTX 3090s, ThinkPad ECU and a VW Golf Radiator Serve 27B Models
TooManyPascals · reddit · 2026-09-26
A Redditor unveiled the "2400cc Inference Racer", a home inference rig doubling as a winter heater:
- Hardware: two used AORUS RTX 3090 XTREME WATERFORCE cards linked via NVLink, hosted by a naked Lenovo ThinkPad motherboard (Ryzen 7 7840U, 64GB RAM) mounted with a custom water block and a clamp.
- Cooling: a €24 VW Golf radiator with mismatched hoses; 2.4L of coolant (hence 2400cc), leaking under 100ml/day.
- PCIe: one GPU on the WWAN slot (Gen4 x1), one on the SSD slot (Gen2 x4); BIOS patched to remove whitelist, alter power-up sequence and disable PCIe power-saving.
- Workload: Ubuntu + vLLM serving Qwen3 27B at surprisingly decent speeds; without the 3090s it acts as a low-idle server running small models via Vulkan/llama.cpp. No hot-swap yet.
More from Infra
- US AI and data center investment to hit 3.63% of GDP yearly, topping the 1800s railroad boom — McDonaghMatthew · 2026-09-26
- Paper: AI buildout to cost $10.3T over 8 years, needs $3.7T annual revenue by 2032 — AccBalanced · 2026-09-26
- Starlink to double maritime speeds to near 1 Gbps by end of December, no new dish needed — XFreeze · 2026-09-26
- News site logs show AI bots hitting it hundreds of thousands of times a day — PJZNY · 2026-09-26
- Why Anthropic lags OpenAI on math: compute constraints and 400k GPUs coming online — haider1 · 2026-09-26
- Unverified: OpenAI halted training of its top upcoming model after Sept 20 incident — ccerrato147 · 2026-09-26