Ling 3.0 Flash on Strix Halo: Beats Qwen in Speed, but Tool Calling is Broken
Badger-Purple · reddit · 2026-08-11
A developer tested the 4-bit compressed Ling 3.0 Flash model on AMD Strix Halo using vLLM (ROCm/HIP). Benchmarks show its inference speed surprisingly outperforms the highly optimized Qwen-122b (rocmFP4). However, the author noted a functional defect: the tool calling feature is currently broken. The post asks the community if others have experienced similar issues.
More from Models
- Hugging Face model update: better multi-character support — OdinLovis · 2026-08-11
- Exploring Knowledge Cutoffs and Pre-training Timelines for Claude & GPT — sshh12 · 2026-08-11
- Grok 4.6 rolling out in Cursor now — daniel_mac8 · 2026-08-11
- Renaming PDF to 'final_draft' Boosts LLM Review Scores: Evaluation Weirdness — CShorten30 · 2026-08-11
- User Tests New Grok Build: Code 'Laziness' Fixed in Minecraft House Task — Angaisb_ · 2026-08-11
- Hands-on with Grok in Cursor: Fast and Agentic, but Lacks Creativity — andrew_n_carr · 2026-08-11