Groq CTO on NUMA Architecture: Per-Core Dedicated HBM Trades Performance for Flexibility

BenBajarin · x · 2026-08-26

This describes a "NUMA-style" memory slicing architecture. Each core gets a view of its local HBM (high-bandwidth memory) slice, eliminating contention in the global memory subsystem, but requiring developers to think carefully about how workloads map onto the machine. It is a trade-off between performance optimization and flexibility.

Original post →

More from Infra

Infra channel →