Running GLM 5.3 Flash NVFP4 inside ChatGPT: a local AI experiment
TheZachMueller · x · 2026-09-20
Developer TheZachMueller shows GLM 5.3 Flash NVFP4 running inside the ChatGPT client, calling it part of his renewed local AI experiments. He says it took significant effort and he plans to write a TIL blog post today documenting how he did it.
An atypical local-deployment combo — details pending in the blog, but it's confirmed to work.
Related event: Dev Runs GLM 5.3 Flash NVFP4 Quantization Inside ChatGPT UI(3 posts)→
More from Infra
- François Fleuret proposes 'AI Safety Levels' air-gapped facilities modeled on bio safety levels — francoisfleuret · 2026-09-20
- Hyperscalers' off-balance-sheet AI commitments top $3.1tn, Morgan Stanley tallies — luisdans · 2026-09-20
- Expanso's Split Architecture: Deterministic Pipelines With an LLM Judgment Layer — aronchick · 2026-09-20
- Jevons paradox is classic low-end disruption you won't spot from a GPU-rich hyperlab — cramforce · 2026-09-20
- SpaceX's orbital AI data centers weigh up to 4,000 kg each, filing seeks 1M satellites — XFreeze · 2026-09-20
- Local Models for Personal Agents: GPT Luna Surprises a Coding-Agent Veteran — gized00 · 2026-09-20