On August 26, 2026, Zhipu’s GLM-5.3-Flash topped overseas testing platforms, securing 62 trillion token calls within a week. The computing foundation for this massive data torrent is powered by a cluster of 100,000 entirely domestic chips. This marks a substantial de-Americanization of China’s AI inference side, two years after the supply of high-end chips was cut off. However, the underlying cost is equally clear: faced with a reality where single-card performance lags by about 4 years, the industry is forcing the technological gap closed by using 10 times the number of chips and exponentially higher electricity bills.
The Hardcore Foundation Behind Anonymous Testing
Before its release, GLM-5.3-Flash, under the pseudonym Ox-Alpha, underwent anonymous testing on two major overseas platforms and quickly dominated the charts. It employs a hybrid architecture with 320B total parameters and 18B active parameters, offering an input price as low as $0.15 per million tokens. This highly aggressive pricing indicates that domestic computing power is not only functional but can also sustain commercial competition in its cost structure.
Figure: The first screen of the GLM-5.3-Flash official release page: 320B total parameters, 18B active, native multimodal. Source: Zhipu official release page screenshot
During the week when overseas users were frantically calling the API, all requests flowed to a domestic chip cluster connected by a self-developed high-bandwidth network. Zhipu adopted a production-grade EPD disaggregated architecture, improving end-to-end service performance by 3 times compared to the baseline. This shows that the domestic computing cluster has surpassed the basic threshold of just being able to run, and now possesses the capability to support production-grade, high-concurrency requests for world-class models.
The Physical Cost of Trading Quantity for Quality
Stripping away the glamorous usage statistics, the physical limitations of the underlying hardware remain harsh. Constrained by the lack of EUV lithography machines and HBM memory export restrictions, mainstream domestic chips still face severe thermal efficiency wall challenges in process technology and memory access capabilities. A rough estimate suggests that the current single-card compute of mainstream domestic chips is only equivalent to about 60% of NVIDIA’s H100 from 4 years ago.
Figure: GLM-5.3-Flash architecture diagram: Linear attention plus sparse attention, reducing KV Cache by 4.44 times under 1M context. Source: Zhipu official architecture diagram
Faced with the latest international cutting-edge architectures, the single-card performance gap can even reach an order of magnitude. In response to this generational gap, Chinese engineers have provided an extremely brute-force solution: relying on highly complex system engineering to smash out total compute with a chip cluster 10 times larger.
Exorbitant Electricity Bills
Piling up massive compute with 100,000 chips inevitably brings an exponential surge in energy consumption. By adopting this quantity-for-compute approach, the power consumption per token could be around 5 times that of the leading international level. The proportion of electricity costs in the cluster’s total cost has skyrocketed from the usual 10% to 20%, to nearly half.
Some overseas developers have also reported that the response speed of Chinese APIs occasionally slightly lags behind top Western providers. Yet, subsidized by massive power generation capacity and an abundant engineer dividend, China’s AI industry has brute-forced a de-Americanized survival path using a high-energy, asset-heavy solution.
An Extreme Test of True Compute
With 62 trillion calls, GLM-5.3-Flash has proven the true availability of domestic chip clusters under massive concurrency. 100,000 chips and soaring electricity bills are the clear price China’s AI is paying to break the technology blockade. When there is an objective generational gap in single-card performance that can only be compensated by ten times the quantity and system-level optimization, this contest has evolved into a battle of system engineering capabilities and strategic resolve.
This path is very expensive, but it has indeed been proven viable. The above analysis is based on current facts.
References:
- Securities Daily Report
- Securities Times Report
- Kama Notes
- Martin Alderson Analysis
- Lobsters Discussion