Skip to main content
Customer case studies

Training-ground domestic GPU inference cluster

In a domestic AI computing ecosystem, compare ZCube with traditional Clos across collective communication, training, and inference.

Training groundDomestic GPU inference clusterIndependent validation

Project information

Project overview

Environment, workload, and validation approach.

Customer type
Training ground
Project stage
Independent validation
Hardware ecosystem
Domestic GPU inference cluster
Workload
Inference and collective communication tests
Comparison baseline
Traditional Clos architecture
Case type
Quantitative validation case

Background and challenges

Project background

Domestic compute clusters need to address congestion and response latency in high-concurrency inference, as well as network hardware spend, and to form a replicable path to large-scale deployment.

Validation or delivery approach

  • Compare ZCube with traditional Clos on the same domestic GPUs
  • Cover collective communication, training, and inference scenarios
  • Independently record All-to-All throughput, TTFT P99, inference throughput, and network spend

Results

Results

The following metrics come from the domestic GPU environment and corresponding test scenarios in this case, and are not merged with other project data.

14.5%

All-to-All throughput increase

Peak throughput advantage in collective communication.

22%–30%

TTFT P99 reduction

Response-latency improvement range in domestic GPU inference tests.

7%–10%

Inference throughput increase

Independent test result for domestic GPU inference.

33%

CapEx reduction

Flat networking reduced network hardware spend.

Value summary

Case value

On domestic-accelerator inference tests, ZCube outperformed traditional Clos on core metrics such as throughput and response latency, providing a technical reference for large-scale domestic compute clusters.

  • Validates ZCube fit for the domestic GPU ecosystem
  • Covers collective communication, training, and inference together
  • Provides independent evidence for large-scale domestic compute cluster construction

Customer case studies

Other customer case studies

See more ZCube validation and production practice across compute environments, workloads, and project stages.

Domestic large-model companyThousand-GPU NVIDIA inference clusterProduction delivery

Large-model company NVIDIA thousand-GPU inference production cluster

Using a coding inference service and comparing against traditional ROFT, this case validates that ZCube unlocks more effective inference compute through a flat fabric and more balanced network communication.

15%

Average GPU inference throughput increase

40.6%

Time to first token (TTFT) P99 reduction

33%

Network hardware cost reduction

View case study
A telecom operatorDomestic GPU clusterIndependent validation

Operator domestic GPU cluster

Compare ZCube with Clos and validate performance gains in collective communication, model training, and multi-tenant end-to-end inference.

119%

Collective communication performance increase

20.1%

Model-training throughput increase

11.5%

Inference latency reduction

View case study
Domestic accelerator companyDomestic thousand-GPU clusterProduction delivery

Domestic accelerator company thousand-GPU production cluster

A domestic thousand-GPU cluster has been live and stable for months, continuously carrying real traffic and serving real users.

Thousand-GPU scale

Domestic compute cluster deployed

Months

Stable production operation

Real users

Continuous service

View case study
Start with the network

Make the next AI data center more efficient, starting with the network

Whether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.

Contact us