Ling 3.0 Flash on Strix Halo: Beats Qwen in Speed, but Tool Calling is Broken

Badger-Purple · reddit · 2026-08-11

A developer tested the 4-bit compressed Ling 3.0 Flash model on AMD Strix Halo using vLLM (ROCm/HIP). Benchmarks show its inference speed surprisingly outperforms the highly optimized Qwen-122b (rocmFP4). However, the author noted a functional defect: the tool calling feature is currently broken. The post asks the community if others have experienced similar issues.

Original post →

More from Models

Models channel →