Skip to main content
Customer case studies

Large-model company NVIDIA thousand-GPU inference production cluster

Using a coding inference service and comparing against traditional ROFT, this case validates that ZCube unlocks more effective inference compute through a flat fabric and more balanced network communication.

Domestic large-model companyThousand-GPU NVIDIA inference clusterProduction delivery

Project information

Project overview

Environment, workload, and validation approach.

Customer type
Domestic large-model company
Project stage
Production delivery
Hardware ecosystem
Thousand-GPU NVIDIA inference cluster
Workload
Coding inference service
Comparison baseline
Traditional ROFT network architecture
Case type
Quantitative validation case

Background and challenges

Project background

Dynamic traffic in large-scale inference easily causes congestion, tail latency, and GPU idle time. The project delivered ZCube at scale on a thousand-GPU production cluster and observed how network architecture optimization affected inference output and unit token cost in real traffic.

Validation or delivery approach

  • Keep GPUs, the software stack, and applications unchanged, and compare production results before and after the network upgrade
  • Adopt ZCube flat networking and more balanced network communication
  • Continuously observe inference throughput, TTFT, KV Cache transfer latency, PFC, and other key metrics

Results

Results

The following metrics apply only to the thousand-GPU cluster, coding inference workload, and network-upgrade conditions described in this case.

15%

Average GPU inference throughput increase

More effective compute unlocked by upgrading the network architecture alone.

40.6%

Time to first token (TTFT) P99 reduction

High-concurrency inference tail latency improved significantly.

33%

Network hardware cost reduction

Required switch and optical-module counts dropped sharply.

27%–46%

KV Cache transfer latency reduction

Improved data exchange between Prefill and Decode nodes.

30%–50%

Network transfer performance increase

Overall results from continued production operation.

79%

Average PFC reduction

Network congestion and PFC backpressure frequency dropped clearly.

Value summary

Case value

On mainstream NVIDIA large-scale inference clusters, ZCube can unlock more effective compute through network architecture optimization and reduce unit token cost.

  • Shows that network architecture innovation can convert directly into production inference gains
  • Unlocks more effective compute without replacing GPUs, the software stack, or applications
  • Provides a reusable practice for network upgrades on thousand-GPU inference clusters

Customer case studies

Other customer case studies

See more ZCube validation and production practice across compute environments, workloads, and project stages.

Training groundDomestic GPU inference clusterIndependent validation

Training-ground domestic GPU inference cluster

In a domestic AI computing ecosystem, compare ZCube with traditional Clos across collective communication, training, and inference.

14.5%

All-to-All throughput increase

22%–30%

TTFT P99 reduction

7%–10%

Inference throughput increase

View case study
A telecom operatorDomestic GPU clusterIndependent validation

Operator domestic GPU cluster

Compare ZCube with Clos and validate performance gains in collective communication, model training, and multi-tenant end-to-end inference.

119%

Collective communication performance increase

20.1%

Model-training throughput increase

11.5%

Inference latency reduction

View case study
Domestic accelerator companyDomestic thousand-GPU clusterProduction delivery

Domestic accelerator company thousand-GPU production cluster

A domestic thousand-GPU cluster has been live and stable for months, continuously carrying real traffic and serving real users.

Thousand-GPU scale

Domestic compute cluster deployed

Months

Stable production operation

Real users

Continuous service

View case study
Start with the network

Make the next AI data center more efficient, starting with the network

Whether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.

Contact us