Running Local Small Models: embeddinggemma Hits Windows Roadblock

QuixiAI · x · 2026-07-21

A developer testing the inference-only embeddinggemma-300M-qat-q40-GGUF (2k context) found that it cannot run natively on Windows and requires WSL. The original poster joked about having picked the best model to save others the hassle, and noted that the model is really fast.

Original post →

More from Infra

Infra channel →