15%
Average GPU inference throughput increase
More effective compute unlocked by upgrading the network architecture alone.
40.6%
Time to first token (TTFT) P99 reduction
High-concurrency inference tail latency improved significantly.
33%
Network hardware cost reduction
Required switch and optical-module counts dropped sharply.
27%–46%
KV Cache transfer latency reduction
Improved data exchange between Prefill and Decode nodes.
30%–50%
Network transfer performance increase
Overall results from continued production operation.
79%
Average PFC reduction
Network congestion and PFC backpressure frequency dropped clearly.