First fully flat fabric
Break the traditional multi-tier tree of data-center networks, remove Spine-layer switches, and flatten the entire interconnect.
Replace the traditional multi-tier tree with a fully flat interconnect, remove Spine-layer switches, and combine single-rail plus multi-rail hybrid access with dedicated routing so GPU communication inside the cluster takes shorter, more balanced paths.

Product positioning
ZCube rebuilds the AI cluster compute network from the bottom up, using less network hardware, shorter forwarding paths, and more balanced traffic to fully unlock cluster compute efficiency.
Core capabilities
Core capability architecture
Physical topology and routing policy are designed together to build shorter, more balanced communication paths for AI workloads.

Break the traditional multi-tier tree of data-center networks, remove Spine-layer switches, and flatten the entire interconnect.
A distinctive NIC access design plans exactly one optimal communication path between any two GPUs in the cluster, improving load balancing structurally.
For the new underlying topology, we built a full set of network control software, including ZCube routing, so hardware topology and software routing work together.
Core hardware and software
The ZCube ecosystem is built from ZCube switches, the agent-native operating system YZOS, the ZCube controller, and supporting deployment tools.
Switch and optical-module counts drop by 1/3, and any two GPUs are 2 switch hops apart; multi-tenant traffic can be physically isolated, and the PFC broadcast domain is optimized.
Built-in AI Agents support congestion-aware adaptive routing and active path switching, plus load-balancing policies based on source and destination IP.
Device state, traffic, microbursts, and port faults are monitored in real time; switches and the controller work together for millisecond-level self-healing.
Supports intent-as-code, automatic construction of domain Skills, validation in real physical environments, and continuous iteration from production feedback.

A core device for high-density AI cluster networking. Specifications follow official product materials.

The ZCube controller covers architecture design, cabling error correction, configuration push, fault monitoring, and diagnosis. It supports automatic recovery from service faults, intelligent RoCE tuning, cluster expansion and hitless replacement, and tenant-level GPU orchestration, forming automated operations, service resilience, and long-term operations capability.
Value and results
Using 51.2T switches (128×400G ports) to build a 16K H200 cluster (NICs at 2×200G) as an example, ZCube improves networking efficiency across device spend, communication path, load balancing, failure domain, and construction time.
Switch count drops from 384 to 256, reducing both network devices and failure sources.
400G module count drops from 49,152 to 32,768.
Fiber cable count drops from 32,768 to 24,576.
Any two GPUs in the fabric pass through only 2 switches, for a shorter communication path.
Architecture and path design together reduce hash collisions and uneven load.
No Spine-switch failures; a switch interconnect link failure affects only some GPUs.
Multi-tenant traffic isolation is raised from logical isolation to physical-link isolation.
Compared with a typical 7–12 days for traditional designs, workloads can go live sooner.
Validation
AI cluster deployments that show ZCube in real environments.
Using a coding inference service and comparing against traditional ROFT, this case validates that ZCube unlocks more effective inference compute through a flat fabric and more balanced network communication.
15%
Average GPU inference throughput increase
40.6%
Time to first token (TTFT) P99 reduction
33%
Network hardware cost reduction
In a domestic AI computing ecosystem, compare ZCube with traditional Clos across collective communication, training, and inference.
14.5%
All-to-All throughput increase
22%–30%
TTFT P99 reduction
7%–10%
Inference throughput increase
Compare ZCube with Clos and validate performance gains in collective communication, model training, and multi-tenant end-to-end inference.
119%
Collective communication performance increase
20.1%
Model-training throughput increase
11.5%
Inference latency reduction
A domestic thousand-GPU cluster has been live and stable for months, continuously carrying real traffic and serving real users.
Thousand-GPU scale
Domestic compute cluster deployed
Months
Stable production operation
Real users
Continuous service
Products and services
Networking, tuning, operations, and professional services work together to cover the full lifecycle of AI cluster networking.
AI computing tuning
For large-scale AI workloads, systematically tune training/inference frameworks, communication libraries, network configuration, and host configuration.
Learn about the productAI computing operations
For large-scale training and inference on AI clusters, build an integrated operations platform across compute, network, and storage for faster problem localization and efficient troubleshooting.
Learn about the productProfessional services
One-stop delivery of network devices, compute-network tuning, and operations services to keep AI cluster networks stable, communication efficient, and SLAs met, helping clusters go live faster.
Learn about the productWhether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.