FEATUREDSynced (机器之心) · WeChat· rssZH03:03 · 05·21
→Zhipu deploys ZCube, raising inference throughput 15% on the same GPUs
Zhipu deployed ZCube in a thousand-GPU GLM-5.1 production inference cluster, replacing ROFT while keeping GPUs, software stack, and business code unchanged; throughput rose by over 15%, TTFT P99 fell 40.6%, and switch plus optical module costs dropped by one third.
#Inference-opt#Zhipu#OpenAI#NVIDIA
why featured
Featured · importance 81 · hook + knowledge + resonance
editor take
Zhipu found 15% throughput in the network layer, not model magic. Strong result, but thousand-GPU inference is not proof for every cluster scale.
sharp
ZCube is a useful reminder that inference margin is now hiding in the fabric, not only in GPU SKUs. Zhipu’s reported setup is unusually clean: a thousand-GPU GLM-5.1 production inference cluster, unchanged GPUs, software stack, and business code, with ROFT swapped for ZCube. The claimed gains are 15%+ throughput, 40.6% lower TTFT P99, and one-third lower switch plus optical module cost.
I buy the direction, not the “overturns twenty years of networking” framing. OpenAI just pushed MRC through OCP in May as a protocol-layer answer to congestion; ZCube attacks structural congestion through topology. Those approaches can coexist. The missing part is workload shape: request distribution, context lengths, and KV Cache traffic share are not given. Without that, 15% is a strong production datapoint, not a portable law.
HKR breakdown
hook ✓knowledge ✓resonance ✓