Compare ZCube with Clos and validate performance gains in collective communication, model training, and multi-tenant end-to-end inference.
A telecom operatorDomestic GPU clusterIndependent validation
Project information
Project overview
Environment, workload, and validation approach.
Customer type
A telecom operator
Project stage
Independent validation
Hardware ecosystem
Domestic GPU cluster
Workload
Multi-tenant end-to-end inference
Comparison baseline
Clos architecture
Case type
Quantitative validation case
Background and challenges
Project background
Operator scenarios need to host multi-tenant workloads on domestic GPU clusters while balancing collective communication, model-training throughput, and end-to-end inference latency.
Validation or delivery approach
Use a Clos architecture without rail optimization as the comparison baseline
Run end-to-end inference tests in a domestic GPU multi-tenant environment
Record collective communication, model-training throughput, and inference latency separately
Results
Results
The following metrics all use a Clos architecture without rail optimization as the baseline, and apply only to the domestic GPU and multi-tenant test environment described in this case.
119%
Collective communication performance increase
Compared with Clos without rail optimization.
20.1%
Model-training throughput increase
Comparison result for domestic GPU model training.
11.5%
Inference latency reduction
Latency improvement in multi-tenant end-to-end inference.
33%
CapEx reduction
The flat architecture reduced network hardware spend.
Value summary
Case value
In multi-tenant inference, ZCube showed a clear performance advantage over Clos without rail optimization.
Validates performance isolation and end-to-end gains in multi-tenant scenarios
Covers collective communication, large-model training, and inference together
Provides a reference path for operator domestic compute cluster construction
Customer case studies
Other customer case studies
See more ZCube validation and production practice across compute environments, workloads, and project stages.
Large-model company NVIDIA thousand-GPU inference production cluster
Using a coding inference service and comparing against traditional ROFT, this case validates that ZCube unlocks more effective inference compute through a flat fabric and more balanced network communication.
Make the next AI data center more efficient, starting with the network
Whether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.