Athena engine runs DeepSeek V4 + Qwen3.8 at 262K context on one DGX Spark
solyarisoftware · x · 2026-09-20
- A new inference engine called Athena has shipped for Nvidia GB10 systems, running DeepSeek V4 Flash and Qwen3.8 Flash-Next at 262K context on a single 128GB DGX Spark with 1,000 tps prefill.
- Benchmarks: DeepSeek V4 Flash hits 1,126 tps at 8K prefill and 948 tps at 256K (19.4 tps decode); Qwen3.8 Flash-Next does 1,071 tps prefill and 32.1 tps decode at 256K — prefill barely degrades from 8K to 256K.
- Key features: disk-cached long conversations (a 141,519-token session restores in 2.1s vs 2min20s fresh), speculative decoding, OpenAI/Anthropic-compatible API, tool calling, Qwen image/document input, persistent agent context, Docker install, and model switching without changing clients.
More from Infra
- Large Power Transformer Lead Times Hit 2029, Prices Up 77% Since 2019, Choking AI Data Center Buildout — sahilypatel · 2026-09-20
- Dev benchmarks Bend 2 on M4 Max: parallel kernel 6.4x faster than NumPy — arthurcolle · 2026-09-20
- China's CXMT says new memory-chip platform enters mass production — johnnyApplePRNG · 2026-09-20
- Unverified claim: OpenAI is out of compute — ns123abc · 2026-09-20
- awesome-local-ai: one-command local AI stacks with coding agents, benchmarked on real hardware — julianharris · 2026-09-20
- Google open-sources AX, a Kubernetes-style orchestrator built for agentic workloads — rakyll · 2026-09-20