Hardware Split Strategy for Self-Hosted Pipelines

TheZachMueller · x · 2026-07-11

The author shares an update on their self-hosting setup, noting that they run **Kimi** and **GLM** on separate halves of a single **B200** node's resources using **NVFP4**. Additionally, the Kimi setup includes a speculative decoder. They also mention waiting for some kernel experiments to finish before integrating and locally hosting **Flash** as well.

Related event: Developer Rebuilds Coding Pipeline Around Self-Hosted GLM(2 posts)→

Original post →

More from coding & agent

coding & agent channel →