Framework tuning
Targets: Megatron, DeepSpeed, SGLang, and others. Techniques: hyperparameter optimization, training/inference performance analysis and tuning, KV Cache tuning.
For large-scale AI workloads, systematically tune training/inference frameworks, communication libraries, network configuration, and host configuration.

Product positioning
YCOS coordinates training/inference frameworks, communication libraries, network, and host configuration, analyzes cross-layer bottlenecks, and raises effective cluster compute and workload performance through end-to-end tuning.
Core capabilities
Core capability architecture
Systematic analysis of key software and hardware configuration along the training and inference path of AI clusters.
Targets: Megatron, DeepSpeed, SGLang, and others. Techniques: hyperparameter optimization, training/inference performance analysis and tuning, KV Cache tuning.
Targets: NCCL, DeepEP, Mooncake TE, and others. Techniques: multi-device fused scheduling and collective communication algorithm tuning.
Targets: RoCE, PFC, ECN, network scheduling, and others. Techniques: PFC+ECN tuning and in-depth network performance analysis.
Targets: CPU, NIC, PCIe, NUMA, and others. Techniques: compute-core scheduling, host-network tuning, and affinity tuning.
Value and results
End-to-end tuning raises compute performance and workload efficiency, and adapts to multiple cluster scenarios.
Increase in total large-model inference throughput.
Reduction in large-model inference TTFT.
Supports multiple inference frameworks, model architectures, and accelerators.
Products and services
Networking, tuning, operations, and professional services work together to cover the full lifecycle of AI cluster networking.
AI cluster networking
Replace the traditional multi-tier tree with a fully flat interconnect, remove Spine-layer switches, and combine single-rail plus multi-rail hybrid access with dedicated routing so GPU communication inside the cluster takes shorter, more balanced paths.
Learn about the productAI computing operations
For large-scale training and inference on AI clusters, build an integrated operations platform across compute, network, and storage for faster problem localization and efficient troubleshooting.
Learn about the productProfessional services
One-stop delivery of network devices, compute-network tuning, and operations services to keep AI cluster networks stable, communication efficient, and SLAs met, helping clusters go live faster.
Learn about the productWhether you are building a new cluster, upgrading an existing network, tuning performance, or deploying intelligent operations, we can assess how the network can unlock more effective compute.