3.6B TwIL-LM3-Pro runs locally on 4GB VRAM, claims 35% lead over VibeThinker-3B

glenbeer · x · 2026-10-04

webAI released TwIL-LM3-Pro, a 3.6B-parameter open-source model built for local inference: the Q4 build is just 2.09 GiB and runs on CPU or 4GB of VRAM, with no cloud or per-query API costs.

Vendor-reported benchmarks:

The release echoes Karpathy's thesis that the next big step in AI comes from tiny models packing dense intelligence. Note all scores are self-reported evaluations.

Original post →

More from Models

Models channel →