openjev-sglang: SGLang Radix Cache lets Qwen3.6-35B-A3B run 64 Jev decisions in under 1s
multiply_matrix · x · 2026-09-25
The SGLang team highlighted Jev-style inference as a perfect fit for its prefix caching and structured generation. ekzhang1's openjev-sglang ships a Jev-compatible public API running the open Qwen3.6-35B-A3B model, using SGLang's Radix Cache to reuse a single shared prefill and finish 64 tasks in under one second.
More from Infra
- Oracle's Massive 'Project Jupiter' Data Center Declares Force Majeure, Jeopardizing US AI Rollout — AIFlow_ML · 2026-09-25
- Wasmer runs a real PostgreSQL 18.4 server on iOS and in the browser via WebAssembly — jedisct1 · 2026-09-25
- Inference startup Jatevo returns: 124B tokens, 1.66M requests, $147K of inference delivered — toptickcrypto · 2026-09-25
- Calibration-Free Quantization Method TQ Open-Sourced, Hits 92.4% Top-1 on Qwen 27B 4-bit — textclf · 2026-09-25
- Agentic AI changes the CPU-to-GPU ratio: 5% CPU allocation cuts token cost ~3.7% — BenBajarin · 2026-09-25
- PreFT Paper Accepted at NeurIPS: Prefill-Only LoRA Adapters Speed Up Multi-Adapter Serving — aryaman2020 · 2026-09-25