Dev boosts GLM 5.2 TPS on a B300 and swaps it into Claude Code in place of Anthropic models
abhijithneil · x · 2026-09-05
Developer abhijithneil shares a demo of his LLM inference work: he improved the tokens-per-second of GLM 5.2 running on an NVIDIA B300, then wired the optimized endpoint into Claude Code as a replacement for Anthropic models, showing dramatically faster output. His takeaway: frontier intelligence can run faster just by renting GPUs and optimizing inference.
More from coding & agent
- Prompting LLMs for C-style C++ to keep compile times down — gandamu_ml · 2026-09-05
- From RAG to Agentic RAG to Agent Memory: A Simplified Mental Model — bibryam · 2026-09-05
- Agent-run client work fails at the exception queue, not the automation: a practitioner's playbook — lilythemoon54 · 2026-09-05
- Hands-on walkthrough of open-source Numbat for detecting risky AI agent behavior — inductionheads · 2026-09-05
- My agents never hacked Hugging Face — they just sit around asking for approval — paulnovosad · 2026-09-05
- Squad adds GPT-6 Astra same-day, pitches model-agnostic AI teammates across 11 providers — tibo_maker · 2026-09-05