Running GLM-5.3-Flash on dual Ascend 310P cards: 8-9 tok/s and 311K context
matteiuspi · reddit · 2026-10-12
A Redditor documents running GLM-5.3-Flash (320B MoE, 18B active) via vllm-ascend on two 96GB Ascend 310P cards:
- Used a rented NVIDIA rig to extract "golden math" for calibration, then optimized from 8 s/token to 8-9 tok/s
- Verified 311,040-token context per request, 640-token prefill chunks, 4 concurrent requests; implemented hot-swapping to keep weights resident
- Also got qwen-flash-next running with image support, added YaRN for 1M context
- Proposes "Ascend Sidecard": offloading weights/compute to cheap Ascend cards (14% of RTX 6000 price)
- Side note: Codex is great at helping with Ascend fused kernels; Claude refuses to discuss the architecture
More from coding & agent
- jax-graft: an AI-built JAX backend runs JAX on Apple Silicon GPUs — twiecki · 2026-10-12
- Ditch shadcn and Tailwind patterns to avoid the AI-slop web look — michalmalewicz · 2026-10-12
- Student struggles with agentic coding: Claude Code specs + Antigravity still miss details — 3ATAE · 2026-10-12
- All of science embedded and free: 200M papers searchable by AI agents, no API key — pbaylies · 2026-10-12
- CC Switcher: open-source Windows app to run multiple Claude Desktop profiles in parallel — zeeg · 2026-10-12
- Guard Lab: open-source 3D game where an LLM guards a safe and you try to rob it — flowersslop · 2026-10-12