Dev boosts GLM 5.2 TPS on a B300 and swaps it into Claude Code in place of Anthropic models

abhijithneil · x · 2026-09-05

Developer abhijithneil shares a demo of his LLM inference work: he improved the tokens-per-second of GLM 5.2 running on an NVIDIA B300, then wired the optimized endpoint into Claude Code as a replacement for Anthropic models, showing dramatically faster output. His takeaway: frontier intelligence can run faster just by renting GPUs and optimizing inference.

Original post →

More from coding & agent

coding & agent channel →