openPangu-2.0-Flash Joins llama.cpp
pmttyji · reddit · 2026-07-19
ikllama.cpp has added support for openPangu-2.0-Flash.
Newly added model specs include:
- 92B-A6B architecture;
- 512K context length;
- Features like MLA-latent cache, DSA/SWA, mHC, and multi-head MTP.
The post also provides a link to the model documentation and the corresponding GGUF download, targeting local inference and llama.cpp users.
More from Models
- Kimi K3 rises to No. 4 on the Agent Arena leaderboard — HeyZoyaKhan · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22