Inkling Open-Source Model Released
BanghuaZ · x · 2026-07-16
SGLang shared and introduced the first open-source model release of Inkling:
- Total parameters 975B, active parameters 41B MoE
- Supports up to 1M context
- Capable of native reasoning across text, image, and audio
- Serving and RL support are simultaneously available, allowing direct execution and fine-tuning on an open-source stack
The post also highlighted its architectural and deployment optimizations:
- The new architecture includes ShortConv, attention with relative positional encoding, and shared expert sink MoE
- Inference-side optimizations include full CUDA graph and MXFP8 KV cache
- Training-side provides full-parameter and LoRA RL, ensuring training/inference consistency through a custom Megatron backend, routing replay, and cross-runtime parameter synchronization
- SGLang and Miles provide Day 0 support
More from Infra
- Should LLM tokens carry green data-center validation labels, like Fair Trade? — jdavid · 2026-09-03
- Commentary: American construction workers want data centers, not just the grey curve — saranormous · 2026-09-03
- Six load forecasters benchmarked on GPU-hours: none beat the last-value baseline — Vegetable-Top-3670 · 2026-09-03
- Visited a 240MW AI data center in Richmond, VA — surprisingly quiet, no high-pitch noise — AndyMasley · 2026-09-03
- Carmack revives rotovators: spinning tethers could slash the cost of space-based data centers — ID_AA_Carmack · 2026-09-03
- GLM-5.3-Flash beats DeepSeek-V4-Flash for writing and vision on 2× DGX Spark — kuhunaxeyive · 2026-09-03