Skip to main content
Products and services

YCOS|AI computing tuning system

For large-scale AI workloads, systematically tune training/inference frameworks, communication libraries, network configuration, and host configuration.

Large-model inference total throughput up 46%Large-model inference TTFT down 22%
YCOS product image

Product positioning

Full-stack tuning from frameworks to the underlying network

YCOS coordinates training/inference frameworks, communication libraries, network, and host configuration, analyzes cross-layer bottlenecks, and raises effective cluster compute and workload performance through end-to-end tuning.

Core capabilities

  • Training/inference frameworks: Megatron, DeepSpeed, SGLang
  • Training/inference communication libraries: NCCL, DeepEP, Mooncake TE
  • Network configuration: RoCE, PFC, ECN, network scheduling
  • Host configuration: CPU, NIC, PCIe, NUMA

Core capability architecture

Coordinated tuning across four key layers

Systematic analysis of key software and hardware configuration along the training and inference path of AI clusters.

Framework tuning

Targets: Megatron, DeepSpeed, SGLang, and others. Techniques: hyperparameter optimization, training/inference performance analysis and tuning, KV Cache tuning.

Communication-library tuning

Targets: NCCL, DeepEP, Mooncake TE, and others. Techniques: multi-device fused scheduling and collective communication algorithm tuning.

Network-configuration tuning

Targets: RoCE, PFC, ECN, network scheduling, and others. Techniques: PFC+ECN tuning and in-depth network performance analysis.

Host-configuration tuning

Targets: CPU, NIC, PCIe, NUMA, and others. Techniques: compute-core scheduling, host-network tuning, and affinity tuning.

Value and results

Full-stack tuning value

End-to-end tuning raises compute performance and workload efficiency, and adapts to multiple cluster scenarios.

46%

Increase in total large-model inference throughput.

22%

Reduction in large-model inference TTFT.

Broad adaptability

Supports multiple inference frameworks, model architectures, and accelerators.

Products and services

Other products and services

Networking, tuning, operations, and professional services work together to cover the full lifecycle of AI cluster networking.

AI cluster networking

ZCubeNext-generation AI cluster networking architecture

Replace the traditional multi-tier tree with a fully flat interconnect, remove Spine-layer switches, and combine single-rail plus multi-rail hybrid access with dedicated routing so GPU communication inside the cluster takes shorter, more balanced paths.

Learn about the product

AI computing operations

XOPSAI computing operations system

For large-scale training and inference on AI clusters, build an integrated operations platform across compute, network, and storage for faster problem localization and efficient troubleshooting.

Learn about the product

Professional services

Full Lifecycle ServicesComplete delivery from planning to go-live

One-stop delivery of network devices, compute-network tuning, and operations services to keep AI cluster networks stable, communication efficient, and SLAs met, helping clusters go live faster.

Learn about the product
Start with the network

Make the next AI data center more efficient, starting with the network

Whether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.

Contact us