Cerebras and Blackwell Power Commercial AFD, Accelerating 3T Param Model Inference 14x

teortaxesTex · x · 2026-08-14

The commercially available AFD (Attention FFN Disaggregation) architecture combines Cerebras and Nvidia Blackwell chips to accelerate ultra-large model inference. The system splits the workload: Blackwell handles Attention and Prefill, while the Cerebras chip executes the Feed Forward Network (FFN).

This setup successfully runs massive models like OpenAI's upcoming 3-trillion-parameter GPT-5.6 Sol, delivering up to 14x the inference speed.

Original post →

More from Infra

Infra channel →