Video summary
美国AI研究员的中国之旅:年轻人,追赶者,算力焦虑与“AGI展示厅” |专访Nathan Lambert【101视频播客】
Main summary
Key takeaways
Tech topics & takeaways (China trip + open-model ecosystem)
China vs. US open-source / “open-weight” model leadership
- Nathan argues the US ceded open-model leadership to China around summer 2025.
- This is based on breakthrough releases from DeepSeek (V3/R1) and Moonshot/Kimi K2, followed by continued competitive momentum.
- On whether the US can catch up:
- He believes Nvidia is the most plausible catalyst, since Nvidia benefits broadly from the success of open models.
- He also notes that major US model labs (e.g., Meta/Microsoft/others) chose not to lean fully into open-weight in recent years.
Best China “open source” models (his current view)
- When pressed, Nathan names Kimi and ZAI as top picks “as of today” for best open models.
- He mentions potential candidates and timelines:
- DeepSeek was the “unanimous king” in 2025, especially the V3 and R1 lines.
- Xiaomi “might” take a crown later.
- He notes that Qwen’s biggest models being closed-source limits adoption/mindshare at the very top sizes.
Why Chinese labs release lots of open models
- Framed as “path to relevance”:
- Open releases help ensure developers/users try the models.
- Otherwise, models can remain effectively invisible even if API access exists.
- Also framed as low-risk relative to their goals:
- The strategy helps build an ecosystem → get feedback → acquire customers.
- Expectation: the “fast follow gap” between labs will continue.
Key constraint: compute scarcity
- Nathan’s main claim: the largest disadvantage for Chinese labs is compute/GPU availability.
- He suggests US labs (Meta/OpenAI) likely have more GPUs than Chinese companies combined.
- He cites a rumor example of individual US researchers having thousands of GPUs.
- Impact on strategy:
- Labs design models around likely downstream usage and training/inference cluster shapes.
- Smaller training clusters may impose a ceiling on scaling to very large parameter frontier models.
- He discusses expectations around roughly 3–6T+ and implies further scaling becomes harder without the latest racks/chips.
- Compute constraints push more work into efficiency.
- He notes that innovations like DeepSeek-style memory saving may spread.
- But US labs may sometimes decide incremental efficiency gains aren’t worth the engineering cost.
Domestic accelerators (Huawei chips)
- Consensus he heard: Huawei accelerators are strong for inference, not training.
- Many labs use Huawei chips for serving and ask Huawei to improve to win more customers.
- He expects this inference/training asymmetry to persist “for years”, with the caveat that outcomes depend on semiconductor progress and ecosystem support for running models.
Model quality gaps: benchmarks vs. real-world utility
- On traditional benchmarks:
- Nathan estimates US models are ~6–9 months ahead.
- But he argues real-world “available performance” can be closer because Chinese labs:
- Release sooner after internal RL/training cycles.
- Example framing: an RL run might take ~a week, then the model can release “within a day.”
- He predicts:
- The benchmark gap may stabilize.
- The US can still pull ahead in intangible areas, such as:
- robustness in “knowledge work”
- early adoption of productized workflows (e.g., cloud coding)
Coding models vs. “ease of use”
- Why a very good Chinese coding model isn’t fully dominant yet:
- The US has an advantage in product UX (e.g., Cursor + Claude/Codex workflows).
- Closed labs get consumer usage data back into training (Cursor/Codex/ChatGPT feedback loops).
- Open-weight labs may produce strong raw code, but usage data often doesn’t feed back as directly.
Open-weight + agent workflows / ecosystem effects
- Agents are an emerging theme:
- agentic products across different price tiers
- knowledge-work tooling
- cheaper experimentation due to open models
- He describes OpenRouter as:
- community-oriented
- smaller than major inference providers (e.g., Fireworks/Together)
- still important for open-model mindshare and co-branding effects.
- Mentions agent examples such as Hermes agent (as part of agent popularity dynamics).
Enterprise strategy differences: why US “doesn’t build open” like China
- He contrasts corporate instincts:
- US consumer instinct: “just use an API key.”
- China corporate instinct: “we have compute/talent; let’s build models and then specialize/fine-tune for internal agentic products,” while still sometimes releasing general models for ecosystem adoption.
- Pattern described:
- Release a general model, then fine-tune internal specialized agents/models
- Example references:
- Meituan-like internal use cases
- Xiaomi as intertwined with robotics/car-cloud interfaces
Alibaba cloud vs. US cloud / distribution model
- Nathan claims China cloud (e.g., Alibaba Cloud) tends to carry its own models.
- By contrast, US ecosystems (e.g., AWS/Bedrock-style patterns) resell/host many models.
- He suggests China lacks some US-style “infrastructure reseller/optimization” layers (e.g., CoreWeave/Nebulous-style equivalents).
Data industry weakness in China (surprising gap)
- He highlights that China’s purchasable external AI training data ecosystem is less developed.
- He contrasts:
- US frontier labs buying data from vendors for evaluation improvements
- China teams more often relying on in-house environments/data
- He argues this can limit top-tier “agentic knowledge work,” which usually requires expensive domain training data (healthcare/law/finance-like labor markets).
Security and alignment framing
- Chinese labs emphasize “security” more than ideological/political “alignment.”
- Nathan supports open models generally but argues diffusion should be paced (e.g., not instantly uploading newly breakthrough closed weights publicly).
Cultural/organizational observations affecting execution speed
- He argues Chinese labs function as fast followers due to:
- heavy young talent pipeline
- fewer “priors” and willingness to adopt new paradigms quickly (MOE, RL, agents, tool-use)
- structured education and competitive gating (which he also suggests may reduce the likelihood of “true visionary” culture at the Ilya Sutskever scale)
- He reiterates compute bottlenecks and organizational coordination:
- building language models is partly an organizational problem
- success depends on breaking big tasks into many parallel contributions.
Named products / systems / model families referenced
- Kimi / Moonshot (including discussion related to “Attention Residuals”; mention of young talent and open footprint)
- ZAI
- DeepSeek (V3, R1; “V3/R1 lines” repeatedly referenced)
- Xiaomi / model releases
- references include a 7B “Mimo model”
- mentions “Flash V2” and “2.5”
- Qwen (open strategy at small-to-medium sizes; big models described as closed)
- Alibaba / ModelScope and broader Alibaba cloud context
- Meituan (agentic products; internal specialized models not necessarily released)
- OpenRouter and agent products such as Hermes agent
- Cursor (plus hosted training/inference workflows, including Cursor training a Kimi model in RL sandboxes and serving via Fireworks)
- Huawei chips (strong for inference; training limitations)
Main speakers / sources
- Primary speaker: Nathan Lambert (AI researcher; interview subject)
- Interview host/source: Valley 101 (host name explicitly cited; exact person not otherwise identified)
- Referenced external figures:
- Dario Amodei (Anthropic CEO; discussed risk framing for open-weight models)
- Jensen (Nvidia; discussed in relation to the China GPU debate)
- Dark matter (referenced as an opponent-side figure in the GPU debate)
- Elon Musk (reposted a Moonshot paper mention)
- Kaifu Lee (01.AI; met during the trip; discussed the leadership/transition question)
- Katherine Rentoul (trip organizer; mentioned once)