Qwen3-Coder 30B runs locally in GitHub Copilot at ~85 tokens/s on 96GB VRAM
ollama · x · 2026-08-13
Developer burkeholland ran Qwen3-Coder 30B locally in the GitHub Copilot app via Ollama, achieving 85 tokens per second on 96GB VRAM. Not as fast as Sol, but promising progress.
More from coding & agent
- Using Codex to Automate Mac Shortcuts for Personal Workflows — billyjhowell · 2026-08-13
- OpenRouter Launches Ori Pi: Run Cloud Models Directly on Raspberry Pi — Scobleizer · 2026-08-13
- Applied Compute Builds Internal AI Agent to Automate Post-Training Workflows — rhythmrg · 2026-08-13
- Agent Model Router Test: 91% Cost Drop, 57% Task Success Rate — kleffew94 · 2026-08-13
- mcp-explorer: The 'curl' for Debugging MCP Servers — JeremyCMorgan · 2026-08-13
- Google shows edge AI on Raspberry Pi with LiteRT and Gemma for real-time tasks — rseroter · 2026-08-13