Running Local Small Models: embeddinggemma Hits Windows Roadblock
QuixiAI · x · 2026-07-21
A developer testing the inference-only embeddinggemma-300M-qat-q40-GGUF (2k context) found that it cannot run natively on Windows and requires WSL. The original poster joked about having picked the best model to save others the hassle, and noted that the model is really fast.
More from Infra
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22