Alibaba's RISC-V XuanTie C950 runs Qwen 3.8B at 30 tokens/s
DeltaSqueezer · reddit · 2026-08-19
Alibaba's T-Head RISC-V processor XuanTie C950 runs the Qwen 3.8B model at 30 tokens/s, as discussed on Reddit. The poster quips "who needs GPUs?", showing that mid-size LLM inference on RISC-V without a GPU is becoming practical — a notable data point for edge/low-power inference.
Related event: Alibaba's Xuantie C950 RISC-V CPU Runs Qwen 3.8B at 30 Tokens/s(2 posts)→
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24