Kimi K3 Is Forcing Silicon Valley to Pick Sides
Kimi K3’s release triggered a split in US AI circles over open weights, China’s model pace, and the economics of frontier labs.

Kimi K3 pushed US AI leaders, policymakers, and startups into an open fight over open weights.
Kimi K3 arrived on July 16 with 2.8 trillion parameters, a 1 million-token context window, and an open-weight release plan that will finish on July 27. Within a week, the model had triggered public attacks from Washington, praise from OpenAI’s Greg Brockman, and open support from Nvidia’s Jensen Huang.
The argument is bigger than one model. K3 forced Silicon Valley to confront a hard question: if a Chinese lab can ship frontier-level performance with open weights and lower pricing, what does that do to the business model of closed labs like OpenAI and Anthropic?
| Metric | Kimi K3 | Comparable data point |
|---|---|---|
| Parameters | 2.8 trillion | Largest open-weight model cited in the source |
| Context window | 1 million tokens | Designed for very long-context work |
| API price | $15 per million output tokens | Less than one-third of Fable 5’s price |
| Artificial Analysis score | 57 | Ranked third, behind Fable 5 at 60 and GPT-5.6 Sol at 59 |
| Decode speedup | 6.3x | Claimed gain from Kimi Delta Attention |
| Scaling efficiency | About 2.5x | Versus Kimi K2 |
K3 is a model, but also a signal
Get the latest AI news in your inbox
Weekly picks of model releases, tools, and deep dives — no spam, unsubscribe anytime.
No spam. Unsubscribe at any time.
K3 is the latest model from Moonshot AI, the company behind the Kimi family. The headline specs are hard to ignore: 2.8 trillion parameters, a Mixture-of-Experts design with 896 experts and 16 active per token, plus a 1 million-token context window.

Those numbers matter because they place K3 in a very small club. In the source article’s reading of Artificial Analysis, K3 scored 57 on the intelligence index, good for third place overall. It also hit No. 1 on the Arena leaderboard for frontend coding within 24 hours of release.
The more interesting part is how Moonshot got there. Kimi Delta Attention claims a 6.3x decode speedup for million-token contexts. Attention Residuals let each layer pull from earlier layers more selectively than a standard residual path. Stable LatentMoE and quantization-aware training from the supervised fine-tuning stage helped lift scaling efficiency by about 2.5x versus K2.
- 2.8 trillion parameters make K3 the largest open-weight model mentioned in the source.
- 896 experts with 16 active per token point to a very large MoE system.
- 1 million-token context is aimed at long documents, codebases, and agent workflows.
- 7.3? No, the source gives 6.3x decode speedup and 2.5x scaling efficiency, both tied to architecture changes.
Washington and Silicon Valley reacted fast
The first public shot came from the White House. On July 22, Michael Kratsios, director of the White House Office of Science and Technology Policy, accused K3 on X of large-scale distillation from US models. Treasury Secretary Scott Bessent also floated the idea of sanctions if Chinese models were found carrying US model watermarks.
That same day, OpenAI president Greg Brockman told Bloomberg that K3 was “a very good model, no question.” He shifted the discussion toward infrastructure, arguing that open-weight models are not free in practice because large-scale deployment still needs expensive hardware. Brockman also estimated that China still trails the US by about four months in overall model capability.
“Open source AI is inherently decelerationist.” — Dean Ball, OpenAI Strategic Future lead, on X
Dean Ball’s post went further. He argued that open models cut into frontier labs’ margins, reduce the money available for infrastructure, and slow down the pace of top-end model development. He even framed a world dominated by open weights as “AI communism,” a phrase that drew immediate backlash.
This is where the split became obvious. Policy hawks want tighter controls on distillation and model access. Closed-lab defenders want to protect the economics that fund giant training runs. Startups and open-model supporters want cheap, usable models that they can actually build on.
Startups and Nvidia pushed back
A new group called Little Tech Association emerged with support from Proton, Replit, and Y Combinator. It helped organize a letter signed by nearly 200 companies warning the White House that banning Chinese open models would hurt US startups first.

Particle founder Suhail Doshi put it bluntly: hundreds of companies could die instantly if they were forced to buy only closed models from providers like Anthropic. That is a practical argument, not an ideological one. For small teams, model price and access decide whether a product ships at all.
Nvidia CEO Jensen Huang took the most surprising position of all. In an Axios interview, he said the US had misread DeepSeek earlier and was misreading K3 now. His view was simple: cheaper AI drives more usage, which increases demand for chips, data centers, and infrastructure.
- Nearly 200 companies signed the startup letter against banning Chinese open models.
- Little Tech Association included Proton, Replit, and Y Combinator.
- Huang argued there is “zero” chance Chinese models push US companies out of the market.
- He also said open models improve safety because more researchers can inspect them.
The economics behind the fight are real
The source article makes a strong case that Ball’s argument is not nonsense, even if it is politically loaded. Frontier labs spend billions to train large models. If a Chinese lab offers a close substitute at a lower price, the profit pool shrinks. Lower profit means less money to reinvest, and public markets may also cut valuations and funding appetite.
That logic explains why closed labs dislike open weights. It also explains why open-model advocates keep pointing to the history of AI itself. Transformer papers were published openly. PyTorch is open source. A lot of the field’s progress came from shared infrastructure before the current wave of closed commercialization.
So the real fight is over the business model. One side wants AI to look like a tightly controlled premium service. The other wants it to look like general-purpose infrastructure, closer to electricity or the internet. Those are different economic systems, not just different products.
The source article also notes a market reaction: the Philadelphia Semiconductor Index fell 12.5% in the week of K3’s release, its biggest drop in 15 months. Nvidia, AMD, and Broadcom all fell, while Chinese AI names such as Zhipu and MiniMax also dropped. That tells you investors were reacting to path conflict, not just benchmark scores.
China’s AI teams are moving on their own clock
Moonshot founder Yang Zhilin has changed course in public view. In 2023, he said closed source was the only path to a super app. After DeepSeek’s shockwave in 2025, Moonshot opened K2, then K2.5, then K3.
That change matters because it shows strategy, not improvisation. Moonshot’s 2026 GTC talk described a plan to replace old assumptions in optimization, attention, and residual connections. The company then moved those ideas into shipping models. Elon Musk called the work impressive, and former OpenAI cofounder Andrej Karpathy said the field may still be underestimating the original Transformer paper.
Moonshot’s own team has also been explicit about constraints. In a Reddit AMA, they said they do not have as many GPUs as US peers, but they squeeze more performance out of each card. Co-founder Zhou Xinyu’s answer to questions about OpenAI’s spending was memorable: “We also don’t know, only Sam knows. We have our own rhythm.”
- DeepSeek, Zhipu, and Moonshot are iterating on different schedules but toward the same frontier.
- Open weights let Chinese labs build reach without matching US capital density.
- Architecture changes matter as much as parameter count in this story.
- K3’s full weights are due on July 27, which will test how fast the ecosystem moves around it.
What K3 changes next
K3 does not prove that open models win. It does prove that open models can force closed labs to defend their economics in public, while also giving startups a cheaper path to build products.
If Brockman is right that China still trails by about four months, that gap is small enough to matter and large enough to keep the race unstable. The next test is not whether K3 gets downloaded, but whether developers fine-tune it, deploy it, and build businesses around it before the month is out.
My read is simple: K3 is less a single release than a stress test for the US AI business model. If the model’s full weights ship on July 27 and adoption spreads the way K2 did, expect more policy noise, more pricing pressure, and more arguments over whether frontier AI should be sold, shared, or somewhere in between.
// Related Articles
- [MODEL]
Google ships Gemini 3.6 Flash and 3.5 Lite
- [MODEL]
Opus 5 lets you ship with fewer refusals
- [MODEL]
Claude Opus 5 undercuts Fable 5 on price
- [MODEL]
OpenAI model catalog adds GPT-5.6 pricing tiers
- [MODEL]
Gemini 3.6 Flash proves Google is betting on efficiency over hype
- [MODEL]
Kimi K3 handles an 820k-line Rust codebase