Artificial Intelligence 3 min read

AI Supernodes Emerge as the New Battleground as China's Tech Giants Race to Build Hundred-Thousand-GPU Clusters

Key Takeaways
  • Frontier models Kimi K3 and GLM-5 both operate at 2.8 trillion parameters and have already been forced to raise prices and cap subscriptions due to compute constraints, with K3 requiring a minimum of 64 GPU cards in supernode configuration.
  • Sugon's Dawn 8000 is the first domestic Chinese system to reach 100,000 GPU cards, while Huawei's Ascend 950 has deployed more than 750 A384 sets at the 1,024-NPU scale.
  • The compute economics are stark β€” industry participants report approximately USD 0.80 in computing costs per USD 1.00 of model revenue, making infrastructure scale the central profit and valuation lever.
  • H3C and Inspur shares surged 58% and 41% respectively within three weeks, and Zhipu AI's stock jumped 37% in a single session after announcing a 1-gigawatt-level data centre build.
  • Industry projections put frontier model parameters at 10 trillion within two years, meaning the supernode infrastructure race is expected to intensify rather than plateau.
Computing power β€” not model sophistication β€” has become the defining constraint in China's AI race, with supernode infrastructure now determining who can compete at the frontier.

Chinese technology companies including Huawei, ZTE, H3C, and Sugon are moving aggressively to build and deploy AI supernode clusters ranging from 1,024 to 100,000 GPU cards, as surging inference demand from frontier models pushes existing computing infrastructure to its limits. The shift has been stark enough that at WAIC 2026, supernodes β€” rather than models themselves β€” emerged as the centrepiece of industry conversation, signalling that computing capacity has overtaken algorithmic capability as the binding constraint on AI progress.

The demand pressure is most visible at the model frontier. Zhipu AI's GLM-5 and Moonshot AI's Kimi K3, both operating at 2.8 trillion parameters, have each been forced to raise prices and restrict subscriptions in response to overwhelming inference demand. The K3 model alone requires 64 or more GPU cards in supernode configuration for deployment, effectively establishing a minimum compute threshold for any organisation seeking to run frontier models. The underlying economics are unforgiving: industry participants report approximately USD 0.80 in computing costs for every USD 1.00 of model revenue β€” but each dollar of revenue corresponds to multiplied valuation multiples, making computing scale the primary growth lever rather than model architecture alone.

Why it matters: the supernode race is rapidly reshaping competitive dynamics across China's AI sector. Companies that cannot secure compute at supernode scale face what analysts in the industry are calling structural exclusion from frontier AI competition. This dynamic is already visible in equity markets β€” H3C shares rose 58% and Inspur climbed 41% within three weeks, reflecting investor recognition that infrastructure providers stand to capture outsized value. Zhipu AI's response to compute pressure β€” acquiring Zhongke Jiahe and announcing plans to build a 1-gigawatt-level domestic AI computing data centre β€” sent its own stock up 37% in a single session, underscoring how capital markets are pricing compute access as a first-order strategic asset.

Technology approaches are diverging significantly among the leading players. Huawei is pursuing a full-stack proprietary model spanning chip to interconnect to cooling, with its Ascend 950 system deploying 1,024 NPUs and more than 750 A384 sets. This approach enables deep cross-layer optimisation but demands substantial ongoing research and development investment. ZTE's OEX platform takes the opposite path, fitting 128 GPUs per cabinet and targeting 10,000-card clusters through an open multi-chip architecture designed to avoid single-chip dependency and give customers greater flexibility. Sugon's Dawn 8000 has become the first domestic system to reach the 100,000-card threshold, while H3C is constructing what it describes as an AI super factory operating at 240 to 300 kilowatts per test point. Smaller players including Moore Threads, Pingtouge, and Baidu Kunlun are targeting cloud and vertical deployment scenarios with a focus on inference efficiency and cost performance. The industry consensus, however, is that no single-point advantage in chip design or software can compensate for gaps in overall system engineering β€” the challenge is maintaining integrity across chip architecture, interconnect, software stack, and operations tooling simultaneously.

The trajectory points firmly toward even greater scale. Model parameters have not peaked at 2.8 trillion; projections from industry participants put the frontier at 10 trillion parameters within two years. This suggests the supernode infrastructure being built today will need to scale further still, and the capital and engineering commitments being made now by Huawei, ZTE, H3C, and Sugon are effectively bets on becoming indispensable to the next generation of AI development β€” not just the current one.

This article was drafted with AI assistance from source reporting, then fact-checked and reviewed by a human editor before publishing. Read our editorial & AI-use policy β†’
Was this article helpful?