Research #08 · Research Report

The Great GPU Migration: Realistic Timelines and Technical Roadblocks in China's Nvidia-to-Ascend Transition

ByYumei Dou
PublishedSeptember 24, 2025
FormatResearch Report
CoverageHardware Migration

Executive Summary

China's strategic pivot from Nvidia to domestic GPUs (Huawei Ascend, Cambricon, Ali T-head) is technically feasible but not frictionless. The transition is neither a binary replacement nor a multi-year impossibility—it's a carefully sequenced, hardware-specific, workload-differentiated strategy where inference and specific training tasks migrate to domestic silicon on 6-12 month timelines, while frontier model training remains Nvidia-dependent for 2-4 years.

The gap between theory (Ascend offers competitive capability) and practice (productionizing Ascend inference at scale) spans three critical friction points: operator/kernel parity gaps (CANN lacks 15-20% of CUDA's niche kernels), device/version matrix complexity (Ascend 910B, 910C variants with different software stacks), and environment stability (production inference requires 99.9% uptime guarantees that Ascend's ecosystem doesn't yet provide).

For operators migrating inference to Ascend, expect 2-3 months for narrow scope (inference-only, supported models, experienced engineering teams) or 6-12 months for broad scope (training + inference, diverse models, less experienced teams). These timelines assume strong organizational discipline and specialized expertise unavailable to most companies.

The strategic outcome by 2028: China maintains dual-stack infrastructure with Nvidia dominating training and serving performance-critical workloads, while Ascend captures 40-60% of commodity inference, serving as the cost-optimized layer of a heterogeneous infrastructure. This isn't replacement—it's specialization.

For investors, the critical insight is that token scaling economics favor inference for the next 5 years. Inference demand grows at 19.2% CAGR (2025-2030, from $106B to $255B hardware equivalent), while training grows at 8.3% CAGR. As inference scales faster than training, the Ascend-appropriate workload share grows automatically, even without organizational migration efforts.

Subscriber Content

Continue reading with a subscription.

The full analysis, data tables, and investment-relevant conclusions are available to Research and Full Quant Intelligence subscribers.