Skip to main content
Products and services

ZCube|Next-generation AI cluster networking architecture

Replace the traditional multi-tier tree with a fully flat interconnect, remove Spine-layer switches, and combine single-rail plus multi-rail hybrid access with dedicated routing so GPU communication inside the cluster takes shorter, more balanced paths.

Inference latency reduced 20%–40%Inference throughput increased 10%–25%Network hardware spend reduced by 1/3Any two GPUs reachable in 2 hops
ZCube product image

Product positioning

Trade hierarchy for efficiency, and determinism for performance

ZCube rebuilds the AI cluster compute network from the bottom up, using less network hardware, shorter forwarding paths, and more balanced traffic to fully unlock cluster compute efficiency.

Core capabilities

  • First fully flat fabric, with Spine-layer switches removed
  • Single-rail plus multi-rail hybrid access to balance AI traffic across switches
  • ZCube dedicated routing and supporting network control software
  • Structural global load balancing and physical multi-tenant isolation

Core capability architecture

From multi-tier stacks to a flat interconnect

Physical topology and routing policy are designed together to build shorter, more balanced communication paths for AI workloads.

Dual-port ZCube 4,096-GPU networking topology

First fully flat fabric

Break the traditional multi-tier tree of data-center networks, remove Spine-layer switches, and flatten the entire interconnect.

Single-rail plus multi-rail hybrid access

A distinctive NIC access design plans exactly one optimal communication path between any two GPUs in the cluster, improving load balancing structurally.

Complete ZCube software stack

For the new underlying topology, we built a full set of network control software, including ZCube routing, so hardware topology and software routing work together.

Core hardware and software

ZCube networking architecture ecosystem

The ZCube ecosystem is built from ZCube switches, the agent-native operating system YZOS, the ZCube controller, and supporting deployment tools.

Efficient networking

Switch and optical-module counts drop by 1/3, and any two GPUs are 2 switch hops apart; multi-tenant traffic can be physically isolated, and the PFC broadcast domain is optimized.

Intelligent load balancing

Built-in AI Agents support congestion-aware adaptive routing and active path switching, plus load-balancing policies based on source and destination IP.

Fast fault recovery

Device state, traffic, microbursts, and port faults are monitored in real time; switches and the controller work together for millisecond-level self-healing.

Agent-native operating system YZOS

Supports intent-as-code, automatic construction of domain Skills, validation in real physical environments, and continuous iteration from production feedback.

ZCube switch product photo

ZCube switch

A core device for high-density AI cluster networking. Specifications follow official product materials.

Switching capacity
102.4 Tbps
Packet forwarding rate
21,000 Mpps
Service ports
128 × QSFP112
Typical power
2,432 W
Chassis size
446 × 800 × 173.6 mm
Port speeds
100G / 200G / 400G
ZCube controller capabilities

ZCube controller: automate and verify large-scale networking

The ZCube controller covers architecture design, cabling error correction, configuration push, fault monitoring, and diagnosis. It supports automatic recovery from service faults, intelligent RoCE tuning, cluster expansion and hitless replacement, and tenant-level GPU orchestration, forming automated operations, service resilience, and long-term operations capability.

Value and results

ZCube versus traditional Clos

Using 51.2T switches (128×400G ports) to build a 16K H200 cluster (NICs at 2×200G) as an example, ZCube improves networking efficiency across device spend, communication path, load balancing, failure domain, and construction time.

33% fewer network devices

Switch count drops from 384 to 256, reducing both network devices and failure sources.

33% fewer 400G modules

400G module count drops from 49,152 to 32,768.

25% fewer fiber cables

Fiber cable count drops from 32,768 to 24,576.

Any two GPUs in 2 hops

Any two GPUs in the fabric pass through only 2 switches, for a shorter communication path.

Structural global load balancing

Architecture and path design together reduce hash collisions and uneven load.

Smaller failure domains

No Spine-switch failures; a switch interconnect link failure affects only some GPUs.

Physical multi-tenant isolation

Multi-tenant traffic isolation is raised from logical isolation to physical-link isolation.

Construction shortened to 5–7 days

Compared with a typical 7–12 days for traditional designs, workloads can go live sooner.

Validation

From deployment validation to production delivery

AI cluster deployments that show ZCube in real environments.

Domestic large-model companyThousand-GPU NVIDIA inference clusterProduction delivery

Large-model company NVIDIA thousand-GPU inference production cluster

Using a coding inference service and comparing against traditional ROFT, this case validates that ZCube unlocks more effective inference compute through a flat fabric and more balanced network communication.

15%

Average GPU inference throughput increase

40.6%

Time to first token (TTFT) P99 reduction

33%

Network hardware cost reduction

View case study
Training groundDomestic GPU inference clusterIndependent validation

Training-ground domestic GPU inference cluster

In a domestic AI computing ecosystem, compare ZCube with traditional Clos across collective communication, training, and inference.

14.5%

All-to-All throughput increase

22%–30%

TTFT P99 reduction

7%–10%

Inference throughput increase

View case study
A telecom operatorDomestic GPU clusterIndependent validation

Operator domestic GPU cluster

Compare ZCube with Clos and validate performance gains in collective communication, model training, and multi-tenant end-to-end inference.

119%

Collective communication performance increase

20.1%

Model-training throughput increase

11.5%

Inference latency reduction

View case study
Domestic accelerator companyDomestic thousand-GPU clusterProduction delivery

Domestic accelerator company thousand-GPU production cluster

A domestic thousand-GPU cluster has been live and stable for months, continuously carrying real traffic and serving real users.

Thousand-GPU scale

Domestic compute cluster deployed

Months

Stable production operation

Real users

Continuous service

View case study

Products and services

Other products and services

Networking, tuning, operations, and professional services work together to cover the full lifecycle of AI cluster networking.

AI computing tuning

YCOSAI computing tuning system

For large-scale AI workloads, systematically tune training/inference frameworks, communication libraries, network configuration, and host configuration.

Learn about the product

AI computing operations

XOPSAI computing operations system

For large-scale training and inference on AI clusters, build an integrated operations platform across compute, network, and storage for faster problem localization and efficient troubleshooting.

Learn about the product

Professional services

Full Lifecycle ServicesComplete delivery from planning to go-live

One-stop delivery of network devices, compute-network tuning, and operations services to keep AI cluster networks stable, communication efficient, and SLAs met, helping clusters go live faster.

Learn about the product
Start with the network

Make the next AI data center more efficient, starting with the network

Whether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.

Contact us