TodayTuesday, September 22, 2026

The Chip Ban Backfired: Huawei Builds a 4,096-Processor AI Cluster

Three years after Washington banned the chips, Huawei unveiled a 4,096-processor AI cluster that the export controls may have built themselves
September 20, 2026
2 mins read
Huawei Atlas 960E SuperPoD AI computing cluster with 4096 Ascend chips unveiled at HUAWEI CONNECT 2026
Huawei unveiled the Atlas 960E SuperPoD at HUAWEI CONNECT 2026 in Shanghai, connecting 4,096 Ascend neural processing units into a single AI computing cluster. [Image Source: Huawei]

SHANGHAI — The United States government spent three years working to ensure that China’s most innovative technology companies could not access the chips they needed to build competitive artificial intelligence infrastructure. Huawei Technologies spent much of that same period building a replacement.

At HUAWEI CONNECT 2026 in Shanghai on Wednesday, Huawei unveiled the Atlas 960E SuperPoD, a computing cluster that connects up to 4,096 of its own Ascend neural processing units into a single unified system. According to Huawei’s announcement, the machine can deliver 8 EFLOPS of FP8 compute performance and 16 EFLOPS of FP4, with up to one petabyte of high-bandwidth memory across the full configuration, sized for training and running models approaching 10 trillion parameters.

The announcement is not just a product launch. It is the most concrete demonstration yet of how US semiconductor export controls, designed to widen the gap between Western and Chinese AI capabilities, have instead become a forcing function for China’s own infrastructure ambitions.

The Atlas 960E introduces Near-Packaged Optics as its core architectural innovation. Where conventional AI clusters rely on large numbers of discrete optical modules to move data between chips, the Atlas 960E replaces those with 5,500 Hi-ONE optical engines embedded directly into the system, reducing complexity while increasing bandwidth. Chips communicate at 7.2 terabits per second using Huawei’s UnifiedBus interconnect. Multiple SuperPoDs can then link into clusters of 512,000 neural processing units; Huawei claims the architecture scales to configurations of one million processors.

From the HUAWEI CONNECT keynote stage, Huawei’s Rotating Chairman David Wang laid out the underlying logic. The company points to what it calls the Tau Scaling Law: as improvements in transistor density slow, gains from system-level interconnect and memory architecture become the dominant driver of AI compute performance. Rather than needing the most advanced chips, the argument goes, you need the smartest cluster design. That argument has particular weight when the most advanced chips are off the table by government fiat.

Washington’s export controls blocked Chinese buyers from accessing Nvidia’s most capable AI processors, including the H100 and its successors. The restrictions were calibrated to prevent Chinese companies from training frontier AI models on domestically available hardware. China’s CXMT announced mass production of its fifth-generation DRAM platform this week, a parallel case of export restrictions generating self-sufficiency rather than dependence.

Huawei Rotating Chairman David Wang delivers keynote at HUAWEI CONNECT 2026 in Shanghai on AI infrastructure strategy
Huawei Rotating Chairman David Wang on stage at HUAWEI CONNECT 2026 in Shanghai, outlining the company’s Tau Scaling Law approach to AI infrastructure. [Image Source: Huawei]
Huawei’s position in the market is harder to read than the product launch suggests. The company says demand for Ascend chips inside China already exceeds available supply, putting a practical ceiling on Atlas 960E deployments until chip production scales. The Ascend 960DT processor, the dedicated training variant that powers the SuperPoD’s full performance envelope, does not ship until the first quarter of 2027.

The software challenge runs deeper than supply. Nvidia’s CUDA platform has been the standard development environment for AI training code for more than fifteen years. It is embedded in virtually every major AI framework and research workflow. Huawei’s CANN software stack exists and supports the Ascend family, but the migration burden for teams running CUDA remains an obstacle that hardware performance alone does not dissolve. No benchmark result accounts for the accumulated tooling and workflow investment CUDA represents.

In the markets where Nvidia can freely operate, the Atlas 960E is not yet a competitive option. But inside China, where entity-list constraints have already demonstrated their limits as a containment strategy, Chinese AI companies including DeepSeek, Baidu, Moonshot AI, and Zhipu AI are training frontier-class models on hardware that US policymakers believed would prove inadequate. The Atlas 960E makes that hardware measurably less inadequate.

The full SuperPoD reaches commercial availability in Q3 2027, with the Ascend roadmap extending to the 970 in 2028 and 980 in 2029. The timetable is longer than an on-stage announcement implies.

What remains genuinely unresolved is the software ecosystem gap. Hardware self-sufficiency is reachable. Huawei is proving that in public, at scale, with each product generation. Whether CANN can deliver the developer experience and framework coverage that would make CUDA migration attractive rather than merely survivable is a different question, and one that a SuperPoD launch does not answer.

Technology Desk

Technology Desk

The Technology Desk leads The Eastern Herald's coverage of consumer technology, online platforms, artificial intelligence, and internet policy.

Leave a Reply

Don't Miss