Local inference hobbyist meme: ChatGPT bros vs Q5_K_M quantized vLLM on CUDA

haydendevs · x · 2026-09-22

An AI-community meme poking fun at the gap between casual ChatGPT users and hardcore local inference hobbyists running NeoHorse-1-9B-Q5KM in vLLM with quantized KV cache on local CUDA.

Original post →

More from Fun

Fun channel →