AMD and Cerebras partner on disaggregated AI inference solution

Go back to AMD and Cerebras partner on disaggregated AI inference solution

AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution

July 23, 2026 1:45 PM EDT

News Highlights 

AMD and Cerebras are collaborating to advance a workload-optimized approach to ultra-low-latency AI inference infrastructure.  AMD Helios and the Cerebras Wafer-Scale Engine will operate as a single disaggregated inference workflow, combining ultra-high-throughput from AMD Instinct GPUs, with ultra-fast token generation of Cerebras Wafer-Scale Engine.Cerebras plans to deploy AMD Helios in its data centers, with the joint solution expected to be available first through Cerebras Cloud in the second half of 2026. 

SAN FRANCISCO and SUNNYVALE, Calif., July 23, 2026 (GLOBE NEWSWIRE) -- AMD (NASDAQ:... More