Inkling Open Model Supports 1M Context
simonguozirui · x · 2026-07-16
The first open Inkling model has been released, highlighting native multimodal reasoning and open-stack support.
Key Details
- Model scale: 975B total / 41B active MoE
- Context length: Up to 1M context
- Capabilities: Natively supports reasoning across text, image, and audio
- Release format: Open weights, supporting serving and RL
Engineering & Deployment Support
- SGLang and Miles provide Day-0 inference support
- New architecture: ShortConv, attention with relative position encoding, shared expert sink MoE
- Deep inference optimizations: prefill full CUDA graph, MXFP8 KV cache
- Training supports full parameter and LoRA RL, ensuring training/inference consistency via a custom Megatron backend, routing replay, and cross-runtime parameter synchronization
Overall, this release feels like a simultaneous launch of both a model and its inference/training stack.
More from Infra
- Reply reiterating: data centers are good for America's construction workers — saranormous · 2026-09-03
- Should LLM tokens carry green data-center validation labels, like Fair Trade? — jdavid · 2026-09-03
- Commentary: American construction workers want data centers, not just the grey curve — saranormous · 2026-09-03
- Six load forecasters benchmarked on GPU-hours: none beat the last-value baseline — Vegetable-Top-3670 · 2026-09-03
- Visited a 240MW AI data center in Richmond, VA — surprisingly quiet, no high-pitch noise — AndyMasley · 2026-09-03
- Carmack revives rotovators: spinning tethers could slash the cost of space-based data centers — ID_AA_Carmack · 2026-09-03