InferCrane: Open-source platform for production-grade deployment and rollback of self-hosted models
yasintoy · reddit · 2026-08-28
InferCrane is open-source infrastructure for operating self-hosted, open-weight, and custom model inference in production. It encapsulates the lifecycle—deployment, routing, scaling, monitoring, and rollback—behind a single stable OpenAI-compatible endpoint. Supporting runtimes like vLLM and SGLang and providers like AWS/GCP/K8s, its core feature, Release Guard, evaluates revisions before promotion. It decides to promote, reject, or mark as insufficient evidence based on operational data, ensuring only safe versions go live, with durable long-running operations.
More from Infra
- Dev runs 125B Qwen model locally on M3 Max at 70 tok/s via MLX — mayfer · 2026-08-28
- NVIDIA launches Mesh open compute network to aggregate idle GPUs for AI — nvidia · 2026-08-28
- AI automated research finds numerical bug in vLLM/SGLang backend — josh_tobin_ · 2026-08-28
- Custom llama.cpp branch speeds up Metal Qwen3.8-Flash-Next inference, adds n-gram SSD offload — tarruda · 2026-08-28
- Hot Chips 2026: inference chips enter an "era of ferment" with divergent bets — BenBajarin · 2026-08-28
- Testing Muon Optimizer: Smoother Gradients and Stable Residual Maxima — stochasticchasm · 2026-08-28