主题综述

World Models:下一个范式? · World Models

主题综述

更新日志

主流共识

第一点:world model 的核心是"学到一个能预测下一状态的可交互表征"

不同团队实现差很远,但这个最底层定义是共享的。General Intuition 的 Pim 讲得最朴素,也点出它和"视频模型"的区别:

"In a video model, you might predict the next likely sequence or the next most entertaining frame. What world models do is they actually have to understand the full range of possibilities and outcomes from the current state, and based on the action that you take, generates the next state, the next frame."
「视频模型预测的是下一个最可能、或最有意思的帧。而 world model 必须理解当前状态下所有可能的结果,然后根据你采取的动作,生成下一个状态、下一帧。」
Pim De Witte (General Intuition) · World Models & General Intuition

Gemini co-lead Oriol Vinyals 给的是最抽象的版本——一种把视频压缩成概念的表征学习:

"A pure aspect of world model would be representation learning. … you could imagine we take these modalities like the videos … And then compressing that into sort of a set of concepts and what those … the movements, the objects, et cetera, are within those … And it models the world in a very compact way that compresses away what's probably not relevant."
「world model 最纯粹的一面是表征学习。……你可以想象我们把视频这些模态……再压缩成一组概念,以及里面的……运动、物体等等是什么……它用一种非常紧凑的方式给世界建模,把大概不相关的东西压掉。」
Oriol Vinyals (Gemini) · Ep 87: Gemini Co-Lead on World Models

第二点:real-time interactivity 是"wow moment"的来源

Genie 团队反复强调即时响应带来的体验跃迁:

"There is something when it responds immediately that is really magical. I think that's kind of sparked the imagination of many people when the DOOM simulation came out …"
「即时响应的时候有一种真正的魔力。DOOM 模拟器出来时,它激发了很多人的想象力……」
"It's at the point where like a human who is not an expert will watch it and think it looks real, right? And I think that's pretty incredible."
「现在到了一个非专家观看时也会觉得'看起来是真的'的程度,对吧?我觉得这相当不可思议。」

第三点:当前都锚定在游戏/视频域——机器人/物理世界是共同的"应许之地"

GI 用游戏剪辑训练、Genie 从游戏引擎演化、xAI 从视频生成切入——三家入口都在游戏/视频,而都把"迁移到机器人和物理世界"当作下一站。Vinyals 把这条说成了 world model 的存在理由:

"It could also meaningfully adds maybe a dimension of simulation that could make us … use, for example, things like prediction before acting in the world. And of course, obvious applications for these kind of 3D or video world models would be clearly … self-driving cars or robotics."
「它可能还能有意义地加上一层'仿真'——让我们能……在真正动手之前先做预测。而这类 3D / 视频 world model 最明显的应用,当然就是……自动驾驶和机器人。」
Oriol Vinyals (Gemini) · Ep 87: Gemini Co-Lead on World Models

分歧在哪

底层定义共享,但再往上分歧很大:world model 到底是什么、智能住在哪一层、以及它部署之后靠什么保持正确。四个阵营之外,还有 Andon Labs 的一盆冷水。

阵营 A · "world model 是概念级表征 / 仿真引擎"——Vinyals 的研究派

Vinyals 的定义最抽象、最雄心——world model 是个"渲染器",你用语言就能改它:

"The world model itself is acting as a renderer of the world that you can really just change by a language."
「world model 本身就像一个世界的渲染器,你用语言就能直接改它。」
Oriol Vinyals (Gemini) · Ep 87: Gemini Co-Lead on World Models

但逐字稿里有一条摘要层完全没收的料——Vinyals 自己都不确定"重力"这种物理概念到底在不在 world model 里,而且语言会"挡在前面",让你没法干净地测它懂不懂物理:

"As soon as you add language, all of a sudden that knowledge is there in the way. So if you ask basic questions about gravity, of course, you would answer them by just having read … explanations of them online and so on. So you would need to somehow connect the concept of gravity — which could be present or not in a world model — to then decode that into an explanation …"
「你一旦加上语言,那部分知识就立刻挡在前面。所以你问关于重力的基本问题,它当然能答——因为它在网上读过……这些解释之类。你需要把'重力'这个概念——它在不在 world model 里都不一定——连接起来,再解码成一个解释……」
Oriol Vinyals (Gemini) · Ep 87: Gemini Co-Lead on World Models

他对机器人迁移也留了余地——还是 open problem:

"For the latter to work better, I mean, it's still a very open problem. There's also all sorts of issues with transfer."
「要让这个(用仿真训练机器人)真正 work,仍然是个很 open 的问题,而且迁移上也有各种各样的问题。」
Oriol Vinyals (Gemini) · Ep 87: Gemini Co-Lead on World Models

阵营 B · "智能在 LLM,视频模型是笨执行器"——xAI 的产品派

xAI 的 Ethan He(曾在 NVIDIA 做 Cosmos world model)抛出一个跟 Vinyals 截然对立的"大胆主张"——视觉智能其实来自语言模型,视频模型本身很笨:

"I have a pretty big claim. The visual intelligence are actually mostly coming from language. … every time you see there, there's some improvement on these models. I would say mostly, again, comes from language model, not coming from the video model itself, like the video distribution models themselves."
「我有一个挺大胆的主张:视觉智能其实大部分来自语言。……每次你看到这些模型有些进步,我会说主要还是来自语言模型,而不是视频模型本身——也就是视频分布模型本身。」
Ethan He (xAI Grok Imagine) · Why Video Agent models are next

他用一个"猫"的例子把"笨"讲得很具体:

"The video distribution models, I would describe they're kind of dumb because they, they take the input instruction literally … If you put a cat in … they would literally show a cat in maybe a white background because you didn't describe the background. The cat is not moving because you didn't describe it. It takes the instruction quite literally. It's kind of dumb."
「视频分布模型,我会说它们有点笨,因为它们把输入指令照字面执行……你输入'一只猫'……它就真给你一只猫,背景可能是白的——因为你没描述背景。猫也不动——因为你没描述它会动。它非常字面地执行指令。挺笨的。」
Ethan He (xAI Grok Imagine) · Why Video Agent models are next

He 据此把未来押到"video agents"——LLM 当大脑、把生成模型当众多工具之一:

"Video agents, mostly language models, will call these generative models … as a tool. So this model can iteratively refine the results or even generate longer content through a very long chain of thought. It's actually very similar to how humans create art. So we don't generate the pixels directly. … It can also use image editing tools from Photoshop."
「video agents 主要是语言模型,把这些生成模型……当成一个工具来调用。这样它就能反复打磨结果、甚至通过很长的思维链生成更长的内容。这其实很像人类创作艺术——我们不直接生成像素。……它也可以用 Photoshop 这类图像编辑工具。」
Ethan He (xAI Grok Imagine) · Why Video Agent models are next

这跟 Vinyals"world model 本身就是那个表征/智能"是直接对立的——一个说智能在 world model 里,一个说 world model 只是手、脑在 LLM。He 还把这套逻辑推到界面层——生成式 UI:

"The generative UI will be user intention to the pixels directly. … I want the email to show to me like a TikTok so I can swipe left and right for the emails. … It's going to be a revolutionary replacement of the interface."
「生成式 UI 就是把用户意图直接映射到像素。……我想让邮件像 TikTok 一样展示,好让我左右滑动切邮件。……它将是界面的革命性替代。」
Ethan He (xAI Grok Imagine) · Why Video Agent models are next

阵营 C · "world model 是从人类行为模仿学来的"——General Intuition 的数据派

GI 的路径既不是 Vinyals 的概念表征、也不是 xAI 的 LLM-driven,而是从海量游戏剪辑做纯模仿学习,直接从帧预测动作:

"What I'm about to show you is a completely vision-based agent that's just seeing pixels and predicting actions the exact same way a human would. … These are pure imitation learning."
「我接下来给你看的,是一个完全基于视觉的 agent——它只看像素、像人一样预测动作。……这些都是纯模仿学习。」
Pim De Witte (General Intuition) · World Models & General Intuition

他把这条路类比成"为交互性造一个 common crawl"——这是 GI 整个赌注的核心隐喻:

"LLMs were trained on predicting like text tokens on words on the internet. What if we predict action tokens on essentially what is the equivalent of the common crawl data set, but for interactivity?"
「LLM 是靠预测文本 token、互联网上的词训练的。那如果我们去预测 action token 呢——在一个相当于 common crawl、但属于'交互性'的数据集上?」
Pim De Witte (General Intuition) · World Models & General Intuition

产品形态上,GI 想直接替换游戏引擎里的"玩家控制器":

"Really what we're doing at the moment is replacing essentially the player controller inside of a game engine. Anything that you're currently … deterministically coding, we hope to replace with a single API, which is just, you stream us frames and we predict actions — and that can be inside an engine or it can be eventually even inside the real world."
「我们现在做的,本质上是替换游戏引擎里的'玩家控制器'。任何你现在……用确定性代码写的行为,我们希望用一个 API 替掉——你把帧流给我们、我们预测动作。这可以在引擎里,最终甚至可以在真实世界里。」
Pim De Witte (General Intuition) · World Models & General Intuition

GI 对这条路的信心强到拒了 OpenAI 给 Medal 游戏剪辑数据开出的 5 亿美元报价、自建独立实验室(Khosla 自 OpenAI 以来最大单笔种子投资)。Pim 的逻辑是数据量足够大、可以并行押模仿学习:

"We essentially realized that we could get so far on just imitation learning. … We think we can essentially leap every single company that's forced to either be consumers of world models or build world models and take this foundation model bet for spatial-temporal agents …"
「我们意识到,光靠模仿学习就能走很远。……我们觉得自己能跳过每一家被迫'要么当 world model 消费者、要么自己造 world model'的公司,直接押这个'时空 agent 的基础模型'赌注……」
Pim De Witte (General Intuition) · World Models & General Intuition

但他对机器人迁移给了一个很诚实的限定——bet 不是迁移到高自由度机器人,而是"机器人得有游戏式输入":

"… the key is that the robot has to have gaming inputs. So our bet is not that we can transfer over to like higher DOF robots and the keyboard and mouse. It's really just that we can move the hard work of pre-training hopefully to post-training."
「……关键前提是:机器人得有'游戏式'的输入。所以我们的赌注不是说能用键鼠迁移到更高自由度的机器人。而只是说,我们能把预训练的苦活挪到后训练。」
Pim De Witte (General Intuition) · World Models & General Intuition

阵营 D · "world model 是持续学习的过程,不是冻结的资产"——Sutton/Oak 的经验派

RL 之父 Rich Sutton 和 Khurram Javed(Oak Lab)从一个前三个阵营都没占的入口切进来:不问 world model 是什么、智能住在哪,而问它部署之后靠什么保持正确。Sutton 把"世界模型"放在 Alberta Plan 的枢纽位置——但前提是持续深度学习:

"There's a very important early step, step two, which is continual deep learning. And we think that one is like almost the most important because it unlocks everything else. If you could do continual deep learning, You could then continually update your model of the world."
「第一步骤是一个非常重要的早期步骤,第二步是持续深度学习。我们认为这一点几乎是最重要的,因为它解锁了其他一切。如果你能进行持续的深度学习,那么你就可以不断更新你的世界模型。」
Rich Sutton (Oak Lab) · Why AI Models Stop Learning

对"人类造仿真器/合成世界"这条路——Genie 的游戏引擎血统、自动驾驶的 sim pipeline、一切人手搭出来的世界模型——Khurram 给出的判据是谁来发现模型错了

"The agent can learn a model from its own experience. And when the agent learns it, it's much better because if the model is incorrect, it can fix it by continual learning. If the humans are making a simulator, then the model only gets updated when the humans figure out that something's wrong. So, yes, planning is important. The agents should learn from simulators, but simulators they make themselves."
「代理可以从自己的经验中学习模型。而当代理学习时,会更好,因为如果模型不正确,它可以通过持续学习来修复。如果人类在制作模拟器,那么模型只有在人类发现某些东西错误时才会更新。所以,是的,规划很重要。代理应该从模拟器中学习,但他们自己制作模拟器。」
Khurram Javed (Oak Lab) · Why AI Models Stop Learning

Sutton 补的一刀更狠——任何人手写出来的世界都是"小世界":

"At first, just say the synthetic data is wrong. I mean, it won't be correct. It'll be a synthetic world. It won't be the real world. And it will matter. The world is incredibly complex. If you write a little program, because there's going to be a little program that will generate the synthetic data, it'll be a small world."
「首先,可以说合成数据是错误的。我的意思是,它不会是正确的。它会是一个合成世界。它不会是真实的世界。这将是一个问题。这个世界极其复杂。如果你写一个小程序,因为会有一个小程序生成合成数据,它会是一个微小的世界。」
Rich Sutton (Oak Lab) · Why AI Models Stop Learning

支撑这套立场的是"大世界假说":世界远比任何 agent 复杂,所以不存在一劳永逸的紧凑世界模型,只有不断重调的严重近似。这直接复杂化了 Vinyals 的"压缩掉大概不相关的东西"(阵营 A)——如果世界大到装不下,压缩出的紧凑表征就必须终身更新,否则从部署那天起就开始过期:

"The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agents, any agent. And this is obvious because the world contains many other agents. … There's no way you can do anything that might claim to be optimal or perfect. You're going to be imperfect and you have to have approximations and those approximations will be severe."
「大世界假说,简单来说,就是世界的复杂性远超你的思维,超过任何代理,任何代理。这很明显,因为世界中包含许多其他代理。……你无法做任何可能声称是最佳或完美的事情。你会不完美,并且必须有近似值,而这些近似值将是严重的。」
Rich Sutton (Oak Lab) · Why AI Models Stop Learning

最后是那句对全领域的裁定——"学到模型、再用模型做规划",这个组合 Sutton 说至今没有实例:

"The big challenge that we don't see in our field, the ability we don't see in our field yet, is the ability to learn a model and then plan with a model. We can do the math things and we can do AlphaGo because the games, we know the model. We know how the moves work. … I can see there's no instances of learning the model and then planning with the model in our field."
「我们在这个领域中看不到的大挑战,是我们在这个领域中还没有看到的能力,是学习模型然后用模型进行规划的能力。我们可以进行数学运算,我们可以做到 AlphaGo,因为游戏中,我们知道模型。我们知道走法是如何工作的。……我看到在我们这个领域没有学习模型然后使用模型进行规划的实例。」
Rich Sutton (Oak Lab) · Why AI Models Stop Learning

按这个标准,三个阵营都还不及格:Genie 生成世界但没人在里面规划,GI 从帧直接出动作(policy,绕开了规划),Vinyals 把"先预测再行动"当应许而非现状。值得注意的是,Sutton 的用法其实是 "world model" 这个词的 RL 原义——agent 自己学、自己维护、拿来做规划的内部模型——比三个阵营的用法都老。他的批评是范式层面的(Oak 自己还没有实证结果),但和 Andon 的实证冷水从两个方向夹住了同一个软肋:当前的 world model 既没证明理解结构,也没证明能被用来规划。

冷水 · "当前模型根本没有空间智能"——Andon Labs 的实证

前三个阵营都假设 world models 在通往物理世界。Andon Labs 把模型放进真实空间推理任务里,撞到的是相反的事实:

"We took models and then we gave them 20 images of interior photographs of apartments. And then we asked them to like redesign the floor plan from that. … you need to reason about 3D space. And it turns out the models are absolutely horrible at this. No one scores statistically better than random chance."
「我们拿了一些模型,给它们 20 张公寓室内照片,让它们据此重新设计平面图。……你需要对 3D 空间做推理。结果发现模型在这件事上烂透了——没有一个的得分在统计上好过随机瞎猜。」
Axel Backlund (Andon Labs) · Reality: The Final Eval

这条直接戳到"world model 通往机器人/物理仿真"的软肋——连 2D→3D 重建都做不好,"物理世界仿真引擎"的应许还很远。而且它和 Vinyals 自己的"我不确定重力概念在不在模型里"是互相印证的。

都没说透的

我的看法

判断(不是事实):当前叫"world model"的至少是四个不同的东西被同一个词收编了——Vinyals 的概念表征(research bet)、xAI 的实时视频交互(product bet)、GI 的行为模仿(data bet),以及 Sutton 拿回来的 RL 原义:agent 自己学、持续更新、拿来做规划的内部模型(paradigm bet)。短期(12–18 个月)跑出商业价值的会是 xAI 那条(视频 / 生成式 UI,最接近现成的消费与创作市场,且务实承认"智能在 LLM");中期最有学术分量的是 Vinyals / GI 那条(如果概念表征或模仿行为真能喂给机器人)。Andon 的"没有空间智能"仍是整条叙事里最该认真对待的信号——如果模型连 2D→3D 重建都不及随机瞎猜,"通往物理世界"就仍是叙事而非路线。Sutton 的批评我买一半:"冻结的 world model 会过期"在逻辑上很难反驳,而且和 Andon 从理论/实验两侧夹住了同一个软肋;但 Oak 自己还没有任何结果,"持续学习解锁一切"目前同样是叙事。我目前的赌注:纯视频/生成路线会先停在创作 / 游戏 / UI 工具;机器人一跳,语料里还没有任何实证支撑。

把握程度:中等偏低。最弱的环节:「生成 ≠ 理解结构」目前只有 Andon 一个实证信源(Vinyals 的"重力不确定在不在"只是旁证),Sutton 的"零实例"是论断不是测量,且各阵营对"智能住在哪"的分歧从未正面碰撞——任何一方拿出一个扎实的机器人迁移结果、或第一个"学到的模型 + 规划"实例,这个判断都可能要翻。

还想知道什么

取材

核心 6 篇按逐字稿核对(2026-06-11 五篇全文重读;2026-08-24 新增 Sutton/Javed 全文通读):

2026-08-10 批次另 3 篇新成员经核查未纳入(Transcript 层零提及 world model,语义擦边):Engram 记忆/持续学习 ×2(38cea6160e7181a498ccefad302369563aaea6160e7181cfbac0e557c83ca7b6)、Kevin Weil AI-for-science(38fea6160e7181d9862cd6a455019494)。

2026-08-24 批次另 4 篇新成员经核查未纳入(Transcript 层零提及 world model,全是 agent/产品层持续学习与私有数据记忆):Trajectory ×3——Arjun Karanam 经验差距演讲(3c4ea6160e7181e0adc7cd856b471297)、Ronak Malde 活系统访谈(3c4ea6160e718148b018d4952c6608e2)、Ronak Malde OPSD 演讲(3c4ea6160e7181d2ba5fef05c86fc589);Jack Morris Engram「Scaling Compute on Context」(3c4ea6160e7181b89f6dc70804705c4f)。

其余 alias 命中、关联较弱或本轮未引用:Project Genie、Fal.ai(生成媒体史)、4D Creation、Best-of-2025 圆桌、John Schulman、Khosla/Rabois Uncapped。后续 headless 全量重综会自动纳入全文。