Sentdex has run 4B+ tokens locally on GLM 5.3 Flash — his most-used local model ever
Sentdex · x · 2026-09-29
ML YouTuber Sentdex reports he's processed over 4 billion tokens locally with GLM 5.3 Flash, calling it by far his most-used local model ever — a strong real-usage endorsement for local inference.
More from Infra
- Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe — DimeRhyme · 2026-09-29
- Redditor predicts sub-$1000 device running SOTA models will spawn the next big company — Robert__Sinclair · 2026-09-29
- Only 3 of ~6,000 data center projects hit by AI buildout moratoriums: SemiAnalysis — MatthewBerman · 2026-09-29
- Cloudflare birthday week ships 8 open source updates: forge, vinext 1.0, native Rust in Workers — ritakozlov · 2026-09-29
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- Kipply breaks down transformer inference arithmetic for H200/B200 in new perf engineering repo — ycombinator · 2026-09-29