The global artificial intelligence hardware landscape is undergoing a profound structural shift as major technology conglomerates race to solve the scaling challenges associated with next-generation large language models. In this high-stakes environment, Huawei Technologies has emerged as a formidable contender, aggressively expanding its domestic hardware ecosystem to navigate ongoing supply chain restrictions and hardware limitations. During a keynote address at HUAWEI CONNECT, David Wang, Deputy Chairman of the Board and Rotating Chairman at Huawei, detailed the company’s comprehensive roadmap for its Ascend processor series and its cutting-edge UnifiedBus interconnect architecture. Designed to handle massive AI models containing up to 10 trillion parameters, these technological advancements represent a strategic pivot toward massively scaled, clustered computing systems as a viable alternative to single-chip brute-force dominance.
The core motivation behind Huawei’s latest infrastructure developments lies in the reality of geopolitical trade constraints. Because international sanctions have restricted Huawei’s access to the most advanced extreme ultraviolet (EUV) lithography equipment and cutting-edge semiconductor fabrication nodes—historically limiting the raw processing power of individual Ascend accelerators relative to top-tier offerings from competitors like Nvidia—the company has chosen to innovate at the system and architectural levels. Instead of relying solely on incremental performance gains within a single silicon die, Huawei is betting heavily on high-density clustering and low-latency interconnect frameworks. By binding thousands of Neural Processing Units (NPUs) into a cohesive, unified computing fabric, Huawei aims to bridge the performance gap and deliver the staggering computational throughput required for the next generation of frontier AI models.
Scaling the Ascend Roadmap: The Atlas 960 SuperPoD Evolution
Central to Huawei’s hardware strategy is the rapid acceleration of its Ascend 960 processor family and its corresponding SuperPoD deployment architectures. According to the updated deployment roadmap presented by leadership, the Atlas 960 SuperPoD configuration is being scaled up dramatically from its predecessor’s framework, which housed 8,192 processors, to an immense system capable of integrating up to 15,488 Ascend 960 processors.
This expanded configuration yields staggering theoretical performance metrics. Huawei reports that a single peak Atlas 960 SuperPoD can deliver an incredible 30 EFLOPS of FP8 computing performance, doubling to 60 EFLOPS when operating in lower-precision FP4 modes. Furthermore, the system accommodates a massive 4,460 terabytes of shared memory capacity, underpinned by an interconnect bandwidth specification of 34 petabytes per second (PB/s). In practical AI training environments, Huawei projects that these configurations will achieve training speeds reaching 15.9 million tokens per second, placing them firmly in the upper echelon of global supercomputing infrastructure.
The development cadence for the Ascend family has also been pulled forward due to unexpected engineering milestones and performance efficiencies exceeding initial internal projections. Huawei confirmed that the Ascend 960DT variant is now officially scheduled for commercial release in the first quarter of 2027, followed swiftly by the Ascend 960PR variant in the third quarter of 2027. Looking further ahead, David Wang outlined a strict, predictable cadence for future hardware iterations, confirming that the Ascend 970 and Ascend 980 chips are slated to roll out in 2028 and 2029, respectively.

"We’re evolving our Ascend chip series on a one-generation-a-year cycle," Wang stated during his keynote address. Highlighting the mathematical foundation of their long-term strategy, he emphasized the integration of the Tau (τ) Scaling Law. "In 2028 and 2029, we will roll out the Ascend 970 and 980 chips, respectively. Thanks to the Tau Scaling Law, not only will their compute specifications continue to double, but you can also expect to see huge improvements across the board."
Unlocking Massive Scale with UnifiedBus and Hi-ONE Optical Technology
To orchestrate thousands of individual processors without incurring catastrophic data bottlenecks, hardware architects must address the formidable challenge of communication overhead. In massive AI training runs, processors spend a significant fraction of their lifecycle exchanging gradient updates and activation tensors across the network fabric. Traditional bus architectures often introduce severe latency and power penalties when scaled beyond a few thousand nodes.
To overcome this hurdle, Huawei has advanced UnifiedBus, an innovative interconnect architecture engineered to link processors, memory modules, and high-speed storage subsystems with minimal communication latency. The company has integrated UnifiedBus directly into the Atlas 960E—a specialized variant of the 960 system capable of tethering up to 4,096 NPUs into a single, tightly coupled domain with uniform memory access across the SuperPoD.
A defining innovation within this architecture is the incorporation of Hi-ONE optical technology, which bridges electronics and photonics by pushing high-speed optical connections deeper into the computing hardware stack. By utilizing near-packaged optics (NPOs), Huawei places optical transceivers into close physical proximity with the computing processors. This design philosophy dramatically shortens the electrical trace lengths required to move data on and off the silicon die, translating directly into higher transmission speeds and substantial reductions in overall energy consumption.
The efficiency dividends of this optical integration are immense. According to technical documentation released alongside the keynote, an Atlas 960E system utilizing Hi-ONE technology requires approximately 5,500 optical units, compared to the staggering 48,000 conventional 800G optical modules typically demanded by legacy architectures. This consolidation yields a direct power savings of over 500 kilowatts specifically dedicated to the processor interconnection network—a critical metric for data center operators grappling with escalating grid power costs and thermal management constraints.
From SuperPoDs to SuperClusters: Building the Million-Processor Horizon

The architectural ambition of the UnifiedBus platform extends well beyond single rack configurations or localized SuperPoDs. Huawei’s long-term deployment matrix envisions the seamless interconnection of multiple Atlas 960E systems into sprawling SuperClusters capable of housing 512,000 NPUs, with an ultimate structural horizon designed to support million-processor configurations.
This scalability is non-negotiable for organizations aiming to train frontier models featuring 10 trillion parameters or more. As neural networks expand in scale and complexity, fitting a model’s weights, optimizer states, and activation memory onto a single accelerator—or even within a standard server node—becomes mathematically impossible. Tensor parallelism, pipeline parallelism, and data parallelism require thousands of computing cores to exchange petabytes of information continuously in real-time. Without a high-bandwidth, low-latency interconnect fabric like UnifiedBus, the computational efficiency of massive clusters collapses under the weight of inter-processor communication delays.
By successfully mating high-density NPU clustering with advanced near-packaged optics and the UnifiedBus protocol, Huawei has formulated a cohesive, end-to-end response to hardware isolation. While international export controls continue to restrict access to premier standalone microprocessors, the company’s aggressive focus on system-level architecture demonstrates that raw silicon performance is only one vector in the modern race for artificial intelligence supremacy.
Broader Industry Implications and Market Outlook
The rapid evolution of Huawei’s Ascend ecosystem carries profound implications for the global semiconductor and AI infrastructure markets. For domestic enterprises within China, the availability of scalable, high-performance domestic AI clusters mitigates corporate reliance on foreign hardware suppliers, creating a self-sustaining ecosystem for local model developers, cloud service providers, and research institutions.
Furthermore, the accelerated one-generation-a-year cadence outlined by Huawei leadership signals to the market that domestic supply chains are overcoming historical manufacturing bottlenecks. By institutionalizing predictable generational leaps up to the planned 2029 Ascend 980 releases, Huawei is providing its enterprise customers with the long-term roadmap visibility necessary to invest heavily in native AI infrastructure.
As the global artificial intelligence community marches toward increasingly massive model architectures and multimodal agentic systems, the competitive battleground is shifting decisively from individual chip specifications to holistic data center engineering. Huawei’s heavy investment in optical interconnects, high-density SuperPoDs, and the UnifiedBus architecture establishes a clear blueprint for how hardware-constrained ecosystems can bypass traditional manufacturing roadblocks through sheer architectural ingenuity. Whether these massive multi-thousand NPU clusters can achieve seamless software ecosystem maturity to rival established global standards remains one of the defining questions for the enterprise technology sector over the remainder of the decade.



