Hardware Split Strategy for Self-Hosted Pipelines

TheZachMueller · x · 2026-07-11

The author shares an update on their self-hosting setup, noting that they run Kimi and GLM on separate halves of a single B200 node's resources using NVFP4. Additionally, the Kimi setup includes a speculative decoder.

They also mention waiting for some kernel experiments to finish before integrating and locally hosting Flash as well.

Related event: Developer Rebuilds Coding Pipeline Around Self-Hosted GLM(2 posts)→

Original post →

More from coding & agent

coding & agent channel →