FuriosaAI Ditches GPU Playbook for 2nm Broadcom-Built Inference Chip, Claims HBM4/E Bandwidth Beats Even the Most Efficient GPUs

TL;DR AI
2 min readKey summary
FuriosaAI and Broadcom are co-developing a third-generation AI inference accelerator built on a 2nm chiplet design.
The chip will use HBM4/E memory and Broadcom packaging, Ethernet, and PCIe technologies for high-bandwidth workloads.
FuriosaAI says the accelerator is optimized for inference, promising better efficiency and token throughput than efficient GPUs.
A compiler-based software stack will map PyTorch models to the hardware, and sampling is planned for the first half of 2028.



