NVIDIA Nemotron 3.5 Lightning Goes Live on Crusoe for High-Volume Agent Inference
Scobleizer · x · 2026-08-11
Crusoe announced that NVIDIA Nemotron 3.5 Lightning is now available on its managed inference platform. The model is a 30B parameter MoE with 3B active parameters, designed for high-volume, specialized tasks in agent workflows.
Crusoe offers flexible deployment options including Serverless and Self-Serve, utilizing MemoryAlloy technology to optimize cluster memory routing and reduce prefill computation.
More from Infra
- Nemotron 3.5 Hits 4,694 tok/s with 64 Concurrent Generations on a Single GH200 — pcuenq · 2026-08-12
- NVIDIA's NeMo Switchyard Cuts Coding Agent Costs by 59% and Runtime by 32% in Ramp Test — sudoraohacker · 2026-08-12
- SD Video Optimization: CK Cuts Generation Time to 473s, but Degrades Prompt Adherence — switch2stock · 2026-08-12
- Muse Glimmer 30B Hits 25 tok/s In-Browser on M4 Max via Custom WebGPU Kernels — xenovatech · 2026-08-12
- Developer asks OpenAI about running ML experiments on basement GPUs — daniel_mac8 · 2026-08-12
- LMSYS introduces Unified Radix Cache: one tree for hybrid model prefix caching — ying11231 · 2026-08-12