Two Years After the Ban: 100K Domestic Chips Power 62 Trillion AI Calls

Two Years After the Ban: 100K Domestic Chips Power 62 Trillion AI Calls

AIDomestic ChipsLLM

Sources:Securities Times + Martin Alderson

On August 26, 2026, Zhipu’s GLM-5.3-Flash topped overseas testing platforms, securing 62 trillion token calls within a week. The computing foundation for this massive data torrent is powered by a cluster of 100,000 entirely domestic chips. This marks a substantial de-Americanization of China’s AI inference side, two years after the supply of high-end chips was cut off. However, the underlying cost is equally clear: faced with a reality where single-card performance lags by about 4 years, the industry is forcing the technological gap closed by using 10 times the number of chips and exponentially higher electricity bills.

The Hardcore Foundation Behind Anonymous Testing

Before its release, GLM-5.3-Flash, under the pseudonym Ox-Alpha, underwent anonymous testing on two major overseas platforms and quickly dominated the charts. It employs a hybrid architecture with 320B total parameters and 18B active parameters, offering an input price as low as $0.15 per million tokens. This highly aggressive pricing indicates that domestic computing power is not only functional but can also sustain commercial competition in its cost structure.

GLM-5.3-Flash Official Release Page Figure: The first screen of the GLM-5.3-Flash official release page: 320B total parameters, 18B active, native multimodal. Source: Zhipu official release page screenshot

During the week when overseas users were frantically calling the API, all requests flowed to a domestic chip cluster connected by a self-developed high-bandwidth network. Zhipu adopted a production-grade EPD disaggregated architecture, improving end-to-end service performance by 3 times compared to the baseline. This shows that the domestic computing cluster has surpassed the basic threshold of just being able to run, and now possesses the capability to support production-grade, high-concurrency requests for world-class models.

The Physical Cost of Trading Quantity for Quality

Stripping away the glamorous usage statistics, the physical limitations of the underlying hardware remain harsh. Constrained by the lack of EUV lithography machines and HBM memory export restrictions, mainstream domestic chips still face severe thermal efficiency wall challenges in process technology and memory access capabilities. A rough estimate suggests that the current single-card compute of mainstream domestic chips is only equivalent to about 60% of NVIDIA’s H100 from 4 years ago.

GLM-5.3-Flash Architecture Diagram Figure: GLM-5.3-Flash architecture diagram: Linear attention plus sparse attention, reducing KV Cache by 4.44 times under 1M context. Source: Zhipu official architecture diagram

Faced with the latest international cutting-edge architectures, the single-card performance gap can even reach an order of magnitude. In response to this generational gap, Chinese engineers have provided an extremely brute-force solution: relying on highly complex system engineering to smash out total compute with a chip cluster 10 times larger.

Exorbitant Electricity Bills

Piling up massive compute with 100,000 chips inevitably brings an exponential surge in energy consumption. By adopting this quantity-for-compute approach, the power consumption per token could be around 5 times that of the leading international level. The proportion of electricity costs in the cluster’s total cost has skyrocketed from the usual 10% to 20%, to nearly half.

Some overseas developers have also reported that the response speed of Chinese APIs occasionally slightly lags behind top Western providers. Yet, subsidized by massive power generation capacity and an abundant engineer dividend, China’s AI industry has brute-forced a de-Americanized survival path using a high-energy, asset-heavy solution.

An Extreme Test of True Compute

With 62 trillion calls, GLM-5.3-Flash has proven the true availability of domestic chip clusters under massive concurrency. 100,000 chips and soaring electricity bills are the clear price China’s AI is paying to break the technology blockade. When there is an objective generational gap in single-card performance that can only be compensated by ten times the quantity and system-level optimization, this contest has evolved into a battle of system engineering capabilities and strategic resolve.

This path is very expensive, but it has indeed been proven viable. The above analysis is based on current facts.

References:

  • Securities Daily Report
  • Securities Times Report
  • Kama Notes
  • Martin Alderson Analysis
  • Lobsters Discussion