llama.cpp Merges Multiple Fixes and MTP Support for Qwen4Exp
jacek2023 · reddit · 2026-09-01
The llama.cpp repository has recently merged several patches targeting the Qwen4Exp (Flash Next) model. Users are advised to update their builds frequently.
Merged PRs (by ServeurpersoCom):
- PR #27978, #28011, #28023, #28123
Merged PR (by 0cc4m):
- PR #28032
In Progress (by danielhanchen):
- PR #27941
Other In Progress:
- MTP support (PR #27836)
- Other fixes (e.g., PR #28136)
More from Infra
- The Next Token Ep 05: AI Inference at Scale, Open Weights, and Industry Burn Rates — threepointone · 2026-09-01
- MongoDB CTO on Database Architecture Evolution and the Unsolved Problem of Agent Memory — The Cognitive Revolution · 2026-09-01
- Stop using long-lived AWS credentials; switch to IAM Roles — _jaydeepkarale · 2026-09-01
- llama.cpp Metal optimization boosts IQ3_XXS decode speed on Apple Silicon — predatar · 2026-09-01
- Architecting pipeline-parallel LLM inference across friends' PCs over the internet — BuildWithEren · 2026-09-01
- Building a Private AI OS on 4x RTX 2080 Ti — askincihan · 2026-09-01