Zyphra Introduces Tensor and Sequence Parallelism (TSP): A Hardware-Aware Training and Inference Strategy That Delivers 2.6x Throughput Over Matched TP+SP Baselines

TL;DR AI
2 min readKey summary
Zyphra unveiled Tensor and Sequence Parallelism (TSP), a new sharding scheme that folds tensor and sequence parallelism onto a single device-mesh axis.
On benchmarks across up to 1,024 AMD MI300X GPUs, TSP lowered per-GPU peak memory versus standard approaches.
It also delivered 2.6x higher throughput than matched tensor-plus-sequence parallel baselines.
The design aims to make large transformer training and inference more efficient, especially for long-context workloads.
