256k Context Inference Hits 100 tok/s
QuixiAI · x · 2026-07-19
Showcases an inference deployment result: running **NVFP4** inference for the **35B-A3B** model at **100 tok/s** using the **QuixiCore SYCL kernel** on **2x Arc Pro B60** GPUs. The focus here isn't on the model itself, but rather that it provides a concrete use case combining **long context (256k) + VRAM/hardware + kernel optimization**, serving as a solid reference for local/edge inference and performance tuning.
Related event: Dual Arc Pro B60 Achieves 100 tok/s on 35B Model Inference(2 posts)→
More from Infra
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Z.AI starts running a 1GW AI data center built entirely on domestic chips — Polymarket · 2026-07-21
- Local models feel far more capable once paired with the right harness — Soft-Barracuda8655 · 2026-07-21