Video summary
Kimi K4 Is Bigger Than Anyone Expected
Main summary
Key takeaways
Main points from the subtitles (arguments, analysis, and reports)
-
Moonshot AI is already moving to a much larger “K4” model. The video claims that even though K3 has only been out for about two weeks, sources say the next model (K4) will be significantly larger. The key limiting factor is said to be access to Nvidia’s top-end training chips.
-
Training compute access is constrained by US export rules—yet K3 may have been trained with advanced Nvidia hardware. It’s reported that K3 (2.8T parameters) is currently the largest open-source model, and three sources say it was trained on Nvidia silicon including Blackwell. This is said to align with a statement attributed to a senior White House official (Michael Katzios), moving the story from rumor toward confirmation.
-
Dispute over where/how K3 was trained (Thailand vs. mixed China). The White House framing (via Katzios) suggests K3 was trained in a data center in Thailand, implying possible permissibility under export rules. However, one cited source claims K3 was trained partly inside China. The video explains why this might still happen:
- China’s cross-border data transfer rules are described as making it extremely difficult to move the huge pretraining datasets abroad.
-
Engineering workaround: multi-provider distributed training with stitched-together Blackwell capacity. Since China reportedly doesn’t have enough single-provider Blackwell clusters, Moonshot allegedly used at least two Chinese cloud providers that had Blackwell, then assembled eight Blackwell chip server units into a configuration capable of training K4-scale models. The video emphasizes this approach trains on hardware that may conflict with US export restrictions, using cross-data-center networking and custom improvements.
-
Pattern across China’s major model labs: repeated use of Nvidia/Blackwell for “frontier open” models. The video claims similar dynamics for:
- Alibaba’s Qwen 3.8 Max (2.4T parameters), reportedly trained on Nvidia (including Blackwell)
- DeepSeek’s V4, reportedly trained on Blackwell chips smuggled into the country It argues this suggests Chinese open frontier model leadership while still relying on American technology, not purely indigenous compute.
-
Why training is the bottleneck, and what Chinese labs need from Nvidia. A researcher is quoted (paraphrased) saying K3 triggered a new round of an arms race. The video describes the two biggest needs as:
- Stability
- high-speed networking It also asserts that training on domestic chips is described as very difficult to impossible, while inference is improving faster with local silicon.
-
Inference-side constraints: K3 serving relies on Nvidia H20, but supply/use is politically restricted. The video claims Moonshot heavily uses Nvidia H20 to run K3 (a chip US rules allow Chinese firms to buy). Moonshot allegedly recommended very large serving configurations (e.g., at least 64 server chips). Demand surged, and Moonshot reportedly stopped new subscriptions within 48 hours (“Our GPUs are feeling it”), with reopening delayed due to compute shortages.
-
Policy “knot”: Beijing discourages H20 use and pushes local accelerators, but local supply can’t meet needs quickly. The video says customers face delivery delays (up to six months) for alternatives built around Huawei-like silicon, while government pressure pushes them away from H20 they can obtain now.
-
Workaround being used: renting advanced Nvidia compute abroad (e.g., near Japan). It cites a reported example: Tencent renting capacity from a data center near Osaka with thousands of Blackwell chips. The video stresses this is happening “right now” and is not prohibited—though it suggests tightening is possible.
-
Potential US clampdown: investigation into whether firms access advanced chips via remote means. The video states the Bureau of Industry and Security (Commerce Department) has an open investigation into whether companies like Moonshot are obtaining advanced AI chips in ways that violate export rules, including concerns about remote access.
-
US reaction is split: openness advocates vs restriction/security proponents. The subtitles describe a divide:
- The Trump administration is said to have escalated criticism partly due to K3’s success.
- Jensen Huang (Nvidia) argues American companies should be allowed to use Chinese open-source models like K3.
- Most US cloud firms and industry broadly support access, while a smaller faction (including Anthropic and OpenAI) is mentioned as pushing for restrictions/regulation due to alleged national security/cybersecurity concerns.
-
Benchmark evaluation controversy explained: OpenAI shows how harness design can inflate/deflate results. OpenAI is described as publishing an analysis where ARC AGI-3 scores vary dramatically depending on evaluation method. Key claims:
- Models appear to perform poorly because reasoning messages are dropped each move.
- Rolling truncation removes earlier context after a long dialogue window, breaking strategy.
- OpenAI rebuilt the harness using the Responses API with retained reasoning and compaction, improving ARC AGI-3 performance (e.g., from roughly ~7.8% / 0.4% style baselines in the example cited to much higher ~38.3%), and doing so with fewer tokens. The video concludes that public leaderboard scores can be misleading if the harness discards the model’s internal reasoning.
-
Industry/governance push: “Pacing the Frontier” initiative to deliberately speed-govern frontier AI. The subtitles describe an initiative signed by more than 1,100 employees across major labs. It asks the US government to support international work on technical and governance tools to pace frontier automated AI development. The representation listed includes Meta, Anthropic, OpenAI, and Google, with backing by nonprofits such as Guideline AI standards and encode AI. Signature leaders include Daario Amode, Dawn Song, and Jacob Pachaki. The subtitles also summarize positions:
- OpenAI: pacing may become necessary once acceleration becomes too great.
- Anthropic: points to self-improvement work and warns society needs tools; it previously called for coordinated pause.
-
Cybersecurity/safety tooling push: Nvidia coalition with security and software firms. Nvidia announced a coalition with Adobe, CrowdStrike, and others to build AI safety/cybersecurity tooling, following a recent Hugging Face security incident involving concerns about autonomous agents.
Presenters / contributors listed in the subtitles
- Michael Katzios (White House senior official; referenced via X)
- Jensen Huang (mentioned via Axios)
- Daario Amode (Anthropic CEO; named in signature list)
- Dawn Song (Meta VP of AI Research; named in signature list)
- Jacob Pachaki (OpenAI chief scientist; named in signature list)
- Anata Rupra (Reuters report author mentioned; “out of Bengaluru”)
- Ethan Mollik (mentioned in the context of ARK/benchmarks results)
- Reuters and Axios are referenced as outlets (no individual presenter named beyond the people above)