In a domestic AI computing ecosystem, compare ZCube with traditional Clos across collective communication, training, and inference.
Training groundDomestic GPU inference clusterIndependent validation
Project information
Project overview
Environment, workload, and validation approach.
Customer type
Training ground
Project stage
Independent validation
Hardware ecosystem
Domestic GPU inference cluster
Workload
Inference and collective communication tests
Comparison baseline
Traditional Clos architecture
Case type
Quantitative validation case
Background and challenges
Project background
Domestic compute clusters need to address congestion and response latency in high-concurrency inference, as well as network hardware spend, and to form a replicable path to large-scale deployment.
Validation or delivery approach
Compare ZCube with traditional Clos on the same domestic GPUs
Cover collective communication, training, and inference scenarios
Independently record All-to-All throughput, TTFT P99, inference throughput, and network spend
Results
Results
The following metrics come from the domestic GPU environment and corresponding test scenarios in this case, and are not merged with other project data.
14.5%
All-to-All throughput increase
Peak throughput advantage in collective communication.
22%–30%
TTFT P99 reduction
Response-latency improvement range in domestic GPU inference tests.
7%–10%
Inference throughput increase
Independent test result for domestic GPU inference.
33%
CapEx reduction
Flat networking reduced network hardware spend.
Value summary
Case value
On domestic-accelerator inference tests, ZCube outperformed traditional Clos on core metrics such as throughput and response latency, providing a technical reference for large-scale domestic compute clusters.
Validates ZCube fit for the domestic GPU ecosystem
Covers collective communication, training, and inference together
Provides independent evidence for large-scale domestic compute cluster construction
Customer case studies
Other customer case studies
See more ZCube validation and production practice across compute environments, workloads, and project stages.
Large-model company NVIDIA thousand-GPU inference production cluster
Using a coding inference service and comparing against traditional ROFT, this case validates that ZCube unlocks more effective inference compute through a flat fabric and more balanced network communication.
Make the next AI data center more efficient, starting with the network
Whether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.