What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use

riceinmybelly · reddit · 2026-09-11

A Reddit user asks whether an 8GB VRAM card like a 2050 is still useful for local inference: loading embeddings, rerankers, and chat models sequentially for office work. He prioritizes tool use and multilingual support over world knowledge, with vision as a nice-to-have.

He feels small models have been neglected lately and wants recommendations for recent good small models.

Original post →

More from Infra

Infra channel →