sparkDash 2.0 open-sources a token-factory dashboard for private GPU fleets
aigclink · x · 2026-10-10
MiaAIlab open-sourced sparkDash 2.0, an operations dashboard for private inference fleets, built on the team's real setup of 4 DGX Spark machines running GLM 5.3 Flash. It treats a GPU cluster like a factory line:
- Fleet status: one card per machine showing GPU, CPU, unified memory, storage, network and the served model; any Linux box with an NVIDIA GPU can join, added from the UI without restart.
- Throughput metrics: reads engine-native stats from vLLM, SGLang, llama.cpp and more — live decode/prefill tok/s, KV cache usage, queues, TTFT, ITL p95, preemptions, prefix cache hits, MTP acceptance.
- Quality gates: built-in Decode bench (1/2/4/8 concurrency), Prefill bench, Quality bench (GSM8K, MMLU, instruction following) and a 69-scenario Tool Eval Bench; results are archived so regressions after swapping models/quantization/engines show up at a glance.
- Cost: token totals by model/day split by generation vs reads, cache hit rates, plus a Fleet energy panel computing power draw, 24h kWh, electricity cost and Wh per 1K tokens — a one-page answer to "is self-hosting cheaper than APIs?".
- Ops: start/stop models from the dashboard, graceful shutdown, Wake-on-LAN and more.
More from Infra
- Starship Could Cut Cost to Orbit 100x to ~$185K Per Ton, Unlocking New Businesses — claud_fuen · 2026-10-10
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- Nvidia CEO's son-in-law becomes VP as Apple cuts iPhone 18 Pro orders 15%+ — 创业邦 · 2026-10-10