Six Months, No Tools: A Chinese AI Catches GPT-5.6

Six Months, No Tools: A Chinese AI Catches GPT-5.6

AIKimiChina AIGPT-5.6Moonshot

Sources:HN + web research · HN

Six Months, No Tools: A Chinese AI Catches GPT-5.6

On July 16, 2026, a Chinese company called Moonshot released a new model named Kimi K3. Two days later, on July 18, Silicon Valley programmer Stephen Bochinski wrote a blog post titled “The Kimi K3 Moment.” The post blew up on Hacker News — 254 points, nearly 300 comments. He put it bluntly in the piece: “In my day-to-day work I run Kimi K3 and Claude side by side, and I can’t tell which is which.”

Not “close.” Not “partially caught up.” Can’t tell them apart.

What makes it sting more for Silicon Valley is the price. K3’s API pricing is a third of Claude’s, and subscriptions start at $19 — while Anthropic (Claude’s maker) has quietly turned off access to its own strongest model at the $20 tier, because it can’t afford to run it.

But what I want to talk about today is a deeper question behind all this: why can a Chinese team catch up in six months to a lead the American teams took years to build?

Kimi K3 article header image Image: Stephen Bochinski’s blog, “The Kimi K3 Moment”

The scorecard: no gear, straight onto the field

Let’s start with the numbers.

On GPQA Diamond, a PhD-level scientific reasoning test, K3 scored 93.5% — the highest of any open model at the time. On the AIME math competition problems and HLE (Humanity’s Last Exam), a brutally hard test bank designed to be “unsolvable for AI,” K3 tied GPT-5.6 without any external tools whatsoever — no search, no calculator, no code execution.

What does that mean? Here’s an analogy: it’s like a car race where everyone else runs with GPS navigation and a pit crew, and K3 shows up empty-handed and posts roughly the same lap time.

Moonshot was candid too, writing directly in its release announcement: “In overall performance, K3 still trails Claude Fable 5 and GPT-5.6 Sol.” But the point is — the gap has shrunk from “hopelessly far behind” to “close enough to put on the same chart.”

From DeepSeek’s first shockwave early in the year to Kimi K3 catching up in July, that’s a total of about six months. In January 2025, the US White House said “the US leads China in AI by roughly 3 to 6 months.” In hindsight, that forecast is unsettlingly accurate.

Kimi K3 benchmark comparison chart Image: Kimi K3’s official technical blog, showing K3’s multi-dimensional comparison against frontier models like GPT-5.6 Sol and Claude Fable 5

Distillation isn’t theft — it’s a law of physics

Which brings us to a word we can’t avoid: “distillation.”

This isn’t the distillation column in a chemical plant. In the AI world, distillation works like this —

Suppose you have a doctoral advisor who knows everything from astronomy to geography, but getting him to teach a class costs a fortune, and he’s getting old and a bit slow to respond. You want to train a young lecturer roughly as capable as the advisor, but cheaper to teach and quicker to respond. What do you do? The smartest approach: have the young lecturer sit in on the advisor’s classes, record every question the advisor answered and every explanation he gave, and then practice repeatedly.

The young lecturer doesn’t need to learn calculus from scratch; he just needs to learn “how the advisor would answer this kind of question.”

That’s distillation. You use a strong model’s (the advisor’s) output data to train a new model (the young lecturer). In theory, the new model can learn the vast majority of the strong model’s capabilities in a very short time.

This has been argued about in the AI industry for a long time. Anthropic calls distillation “industrial-scale theft,” and OpenAI says DeepSeek “clearly distilled our models.” Meanwhile, the top-voted comment on Hacker News put it this way:

“A distillation attack isn’t an attack. American frontier labs ‘distilled’ all of humanity’s writing into their own models; it was always going to end with second-tier labs ‘distilling’ their models into cheaper versions. No one can stop anyone else from saving the chat logs and training their own model. The ending was written from the start. Invest in hardware companies, not model companies.

It sounds harsh, but the logic is hard to rebut. When you trained your model on the whole internet’s text, you didn’t go ask each author for permission. Now someone uses your model’s output to train a new model and you call it theft — that moral high ground isn’t especially firm.

An ever-shallower moat

Distillation is deadly because it digs the leader’s moat shallower.

In traditional industries, what are a leader’s barriers? Patents, brand, supply chains, customer relationships. These don’t vanish just because a competitor took one look at your product.

But an AI model’s “knowledge” is essentially parameters — just numbers. When your model answers a question, it outputs a string of text. That string is a knowledge leak. As long as someone is willing to put in the effort to collect millions of question-answer pairs and feed them to a new model to train on repeatedly, the new model can learn to “answer like you.”

This process needs no stolen source code, no poached engineers, not even especially powerful hardware — just patiently collecting data and then training. And this is exactly where Chinese teams have the advantage: low labor costs, strong engineering execution, and zero constraint from US copyright law.

Some HN users pointed out that in early testing, Kimi K3, when asked “what’s your name,” answered “Claude” — which practically means a large volume of Claude conversation logs was stuffed into the training data.

But the sad part is: even if confirmed, so what? US courts have no reach over Chinese companies, and in a cross-border context an Anthropic lawsuit is essentially wastepaper. As one commenter said: “Chinese labs can freely train on pirated data, while American labs have to pay $1.5 billion to settle class-action suits.”

Who won? Who lost?

Let’s start with the losers.

Model companies, especially those whose moat is built on “my model is the smartest.” If your edge gets matched every six months, where’s your pricing power? Anthropic turning off Fable access at the $20 tier is, in essence, saying: our business model can’t survive at this price. When a product costs more than its selling price, and a competitor is selling roughly the same thing at a third of the price — that business can’t work.

US AI regulation. Bochinski makes a sharp point in the piece: the US government held back its own models from release, layered on safety review after safety review, and the result? A Chinese lab outside US jurisdiction just open-sourced an equivalent model. “The only thing this regulation restricts is American users.”

Now the winners.

Hardware companies. No matter whose model you use, you have to buy GPUs to run it. Nvidia doesn’t care whether the model being trained is American or Chinese — either way the chips are theirs. This is the core logic of that top HN comment.

Ordinary users. If you’re a regular reader on WeChat, you don’t need to care whether distillation is legal or whether the model has 2.8 trillion parameters. All you need to know: where a month used to buy you $20 for a nerfed AI assistant, now $19 gets you something at world-class level — and it’s open source, meaning in theory you can download it and run it yourself, no company required.

This is especially good for Chinese users. Over the past two years, American AI products have grown increasingly unfriendly to Chinese users — registration requires a US phone number, payment requires a US credit card, content moderation keeps tightening. Now the most frontier capability sits in the hands of a Chinese company, with no access restrictions and no censorship nerfing.

Kimi K3 multi-dimensional benchmark radar chart Image: Kimi K3’s official technical blog, showing K3’s comprehensive comparison against frontier models across capability dimensions

What happens next?

I think this story is only just beginning. Kimi K3 won’t be the last Chinese model to catch up.

If distillation really is a viable path, it means any breakthrough US labs make in model capability will be matched within a few months. The lead OpenAI and Anthropic burned billions of dollars to build may only be worth a six-month exclusivity window.

This will push the industry in two directions:

One direction is building applications. Don’t sell the model, sell the ability to solve problems. Like electricity — no one cares which power plant generated it; people care whether the fridge is cold and the TV is bright.

The other direction is building hardware. That HN comment, “invest in hardware companies, not model companies,” is investment advice, but it points to a fact: in the AI supply chain, hardware is the one link you can’t route around.

But what worries me most is a third possible direction: the US government’s reaction. Bochinski closes the piece with a passage roughly to this effect — the US government may well go down the same road as the auto industry: propping up a batch of domestic models with subsidies, tariffs, and industrial protection that can only be used inside the wall, expensive and not good enough. In the end, America becomes the one country that can’t get the best AI.

That forecast sounds like fearmongering, but think about the state of the US auto industry, and it doesn’t seem entirely impossible.


Writing this, I’m reminded of another HN comment:

“American labs scraped everything on the internet to train their models, and now that others ‘distill’ their models, they start crying.”

To be fair, the two things aren’t legally identical. But from an ordinary person’s viewpoint, it does feel a bit alike — you took all of humanity’s knowledge for free, yet you demand that others pay to use yours.

After K3’s release, how Anthropic and OpenAI respond, where US policy turns, whether hardware or models eat first — every one of those questions is worth watching. But one thing is already clear: the lead an AI model holds is thin as a sheet of paper.


Reference links:

  • Stephen Bochinski: The Kimi K3 Moment
  • HN discussion (item?id=48960218)
  • Kimi K3 official technical blog
  • Codersera: Kimi K3 Benchmarks vs Fable 5, GPT-5.6 & Opus

Disclaimer: This article is based on publicly available blog posts, HN community discussion, and public industry data. It reflects only the author’s personal observations and does not constitute investment advice.