主题综述

模型表现触顶了吗 · Performance Plateau

主题综述

更新日志

主流共识

第一点:pre-training 这一条轴的回报在变缓——几乎所有人,包括"看多"派,都接受这一点。

AI Vibe Check 的几位研究者认为:简单地往模型里塞更多数据,回报正在递减;与其无限砸资本,不如聚焦样本效率(sample efficiency)。(该集英文原音、podwise 仅存中文,故转述。)

收益递减是固有的,因为模型的智能与用于训练它的计算量呈对数线性关系,这意味着你必须以指数方式增加计算量才能获得智能的每次增量。
— Bob McGrew (former OpenAI) · 见姊妹主题 inference-economics

2026-07 更新:这条"接近事实"的共识第一次出现了正面异议。Mark Chen(OpenAI 首席研究官)不接受"回报变缓"的叙事框架本身:

"I think pre-training is definitely not dead. It's underrated."
「我认为预训练绝对没有死。它被低估了。」

注意他反对的不是对数线性的数学(访谈里他没有正面处理这个问题),而是"瓶颈会真的挡住 scaling"这个推论——完整论证见阵营 B。

2026-08 更新:这一条正式从"接近事实"降格为"有名有姓的两边对峙"。确认侧来了语料里最硬的资历——Igor Babushkin(DeepMind/OpenAI/xAI 三段经历、xAI 联合创始人,亲手参与过大规模预训练)直说"训练确实有递减收益、我们正在接近那个点"(逐字见阵营 A);异议侧则加入了 Sam Altman 本人("scaling laws 是史上最被讨厌的预测——但它一直在继续",见阵营 B)。同一家公司走出来的人开始站到这条共识的两侧,这本身就是信息。

第二点:capability 的瓶颈正在远离"模型本身",开始向部署/使用侧迁移。即使最看好的模型公司也承认这点。

"99% of people get to use bad tools or don't have any tools at all."
「99% 的人用的是糟糕的工具,或者根本没工具。」
Brad Lightcap (OpenAI) · Uncapped #46 Brad Lightcap from OpenAI
"We're so far from just the ability of the models right now being integrated into daily life. People do not know how to use these systems."
「我们离把模型现有的能力真正整合进日常生活还差得很远。人们根本不知道怎么用这些系统。」
Winston Weinberg (Harvey) · 20VC: How Model Performance is Plateauing

第三点:新增量正在出现在 reasoning / post-training / data-quality 等非"参数 + 算力"的轴上——但各家押不同的子轴。

2026-08 补:这一点有了最简洁的官方表述。Sam Altman 把"plateau 在哪"重述成"瓶颈在轮转"——一个动态框架,直接消解了把任何单一时刻的瓶颈当成终点的读法:

"We just had to scale up. We were only bottlenecked on compute. Then we ran out of data and we were bottlenecked on data. We had to figure out what to do there. Now again, I would say we are still bottlenecked on compute, but the last six months or whatever have been a real triumph of a time for research ideas again. So there's always a bottleneck, but the bottleneck moves around."
「我们只需扩大规模。我们只是在计算资源上遇到了瓶颈。然后我们数据耗尽了,遇到了数据瓶颈。我们必须弄清楚该怎么办。再说一次,我认为我们仍然受到计算能力的限制,但过去六个月或者说这段时间又是研究想法的真正胜利时期。所以总是会有瓶颈,但瓶颈会变化。」

注意这个框架的修辞功能:它让"某条轴变缓"永远不构成对 scaling 叙事的反例——任何减速都可以被重述为"瓶颈移到了别处"。好用,但不可证伪。

分歧在哪

阵营 A · "pre-training 平台期是真的"——Harvey / Vibe Check panel 立场

Winston Weinberg (Harvey) 把"plateauing"明确写进了播客标题——但他的论证更微妙,他强调的不是"capability 没涨",而是部署侧已经赶不上,所以应用层公司的现实约束已经不在模型升级:

"We're so far from just the ability of the models right now being integrated into daily life."
「我们离把模型现有能力整合到日常生活,还差得很远。」

Vibe Check panel 给的是更技术化的"plateau"——但他们也加了重要限定:

他们还有一个更技术化的观察(同为转述):RL 能不能成功,取决于落进一个"适度区"——模型已经懂得够多、能做出合理猜测,但又还没强到能直接解决任务;太难或太易,RL 都吃不到信号。

"We're just starting to scratch the surface in terms of economic value creation from the model."
「就模型创造的经济价值而言,我们才刚刚触到表面。」
Vibe Check panel · AI Vibe Check

注意——这个 panel 是"plateauing"派里最看多经济价值的,他们对"capability 平台"和"应用价值平台"的区分极为锋利

Jack Morris 在更基础的研究层面给了一个具体的"plateau 测得到"的论点:

"So we have this result that's maybe the third discovery I was alluding to, which is a way to measure the exact capacity of a language model. And we get this number, if you train a language model on a ton of random data, and you measure its rate of memorization … No matter how you scale the training size, you hit this perfect perfect-ish plateau in model memorization, which we call the model capacity."
「所以我们有一个结果,这可能是我提到的第三个发现,这是一种测量语言模型确切容量的方法。我们得到了这个数字,如果你在大量的随机数据上训练一个语言模型,然后测量它的记忆率……无论你如何缩放训练规模,你都会在模型记忆中达到一个完美的平台期,我们称之为模型容量。」

Igor Babushkin(xAI 联合创始人,DeepMind/OpenAI 出身,现创办 River AI)(2026-08 新增)是这一阵营到目前为止资历最硬的声音——语料里第一个亲手跑过 frontier 预训练、然后正面说出"递减收益"的人:

"Yeah, so training of these models actually has diminishing returns. So the stronger we want to make the pre-trained model and the model overall, the more GPUs we have to put together, the more high-quality data we have to gather. … But eventually something will have to give. So at some point, you can't cover the entire earth with GPUs."
「是的,这些模型的训练实际上会有递减收益。我们希望使预训练模型和整体模型更强大,就需要更好的 GPU 进行搭建,还需要收集更多高质量的数据。……但最终总会有些事情必须让步。因此在某些时候,你无法用 GPU 覆盖整个地球。」

而且他不是泛泛而谈"终有一天",他给了时间判断,并把结构性后果推到底——专有模型商会被两头挤压:

"So eventually there will be a bit of a slowdown in terms of what the AI companies are able to do in terms of The capabilities of their models. And I would argue we're starting to get close to that point. … At the same time, open models are getting stronger and stronger. So I think as a proprietary model builder, you're kind of starting to get squeezed in."
「所以,最终在人工智能公司在模型能力方面所能做到的事情上,可能会出现一点减速。我认为我们开始接近这个节点。……与此同时,开源模型正变得越来越强大。我认为作为一个专有模型的构建者,你有点开始受到挤压。」
Igor Babushkin · Ep 92: xAI Co-Founder

主持人追问"所以你认为能力提升基本在放缓",他答"Yeah, exactly"——没有留退路。他还补了一刀:集中式后训练(把全世界领域专家的知识汇进一套权重)同样在遇到递减,他给的出路是"去分布式"、让企业和个人在本地做自己的后训练。利益结构照例要记账:这正是他新公司 River AI 卖的东西——"集中式训练触顶"对他而言既是判断也是商业前提,和阵营 F 的 OpenAI 立场是同一枚硬币的镜像。

阵营 B · "Scaling laws will continue"——但要扩展定义

Joelle Pineau (Cohere Chief Scientist) 是这一派最公开的代表,名义上看多 scaling,但她的论证已经把"scaling"重定义到包含算法创新:

"I tend to decompose different ingredients that lead to progress. You know, people often talk about like the algorithms, the data, the compute. I think in general, compute and data have a more linear effect on progress. You build more compute, you run bigger models, you can typically get better performance, you feed in more data."
「我习惯把推动进步的不同要素拆开——大家常说的算法、数据、算力。总体上,算力和数据对进步的影响更线性:你堆更多算力、跑更大的模型,通常性能就更好;你喂更多数据也是。」

——言下之意,算法创新才是那条非线性的轴,只是它要很久才看得出来。这就是她把"scaling 会继续"悄悄重定义成"算法那条轴会继续"的地方。

她对 RL 既看多本质、又给了一个限定——别指望开箱即用的 RL 直接给出 AGI,这是 Camp B 内部值得注意的微差异:

"I'm still super bullish on RL in that the concept itself is so fundamental. This idea of training through a system of rewards of indicating what's valuable and what's not valuable through numerical values, that is so fundamental. It's not going away. Now, you know, where we're maybe getting a little bit ahead is thinking that just RL out of the box is going to give us AGI. That part, a lot less so."
「我对 RL 仍然非常看多——因为这个概念本身太根本了。通过一套奖励系统、用数值指出什么有价值、什么没价值来训练,这个想法太根本了,不会消失。现在,我们也许有点超前的地方,是以为'开箱即用的 RL'就能给我们 AGI——这一点,我就没那么信了。」

Mark Chen (OpenAI Chief Research Officer)(2026-07 新增)是这一派目前最"原教旨"的声音——与 Joelle 恰成对照:Joelle 把"scaling 会继续"重定义为"算法那条轴会继续",Mark Chen 则直接为最狭义的 scaling laws 辩护,论据是十个数量级的历史:

"So I mean, it's held for almost 10 orders of magnitude. There's no reason it should not keep holding."
「所以我的观点是,这已经持续了近 10 个数量级。没有理由它不会继续保持下去。」

他对"pre-training is dead"叙事的处理方式是历史归纳——每一代瓶颈都曾被宣布为终点、又都被突破:

"There have always been some bottlenecks that people, well, you can't scale past this because of this bottleneck. And we've always found some kind of technique, whether it be better engineering or some new research insight that helps you break past the boundary."
「总是有一些瓶颈,人们会说,你不能超过这个瓶颈。而我们总是能够找到某种技术,无论是更好的工程还是新的研究见解,帮助你突破界限。」

值得对照:Joelle 用要素分解论证、Casado 用资本可追溯性论证、Mark Chen 用历史归纳论证——阵营 B 内部三种论证类型,强度和可证伪性各不相同。而 Mark Chen 自己也给 RL 那条新轴画了边界(与 Joelle 的"别指望开箱即用 RL"形成 Camp B 内部第二组呼应):

"So it's these fields where things are hard to grade, where RL has the least amount of ability to go and directly apply there."
「所以这些领域的东西难以评估,强化学习在这里的直接应用能力最弱。」

Sarah Wang / Martin Casado (a16z) 把"scaling continues"绑到资本流动模式上:

"This is probably also a unique time in that for the first time you can actually trace dollars to outcomes … provided that scaling laws are holding and capabilities are actually moving forward."
「这可能也是一个独特时期——你第一次能把美元真正追溯到结果,*前提是 scaling laws 仍然成立、能力仍然在前进*。」
Martin Casado / Sarah Wang · Inside AI's $10B+ Capital Flywheel

注意"前提"两个字——这是 Camp B 里自己埋下的撤退条款:如果 scaling 真停了,整个资本飞轮的逻辑会断。

Sam Altman(2026-08 新增,两集访谈)加入后,Camp B 的历史归纳论证有了最高级别的版本——与 Mark Chen 的"10 个数量级"同型,但更直白地把怀疑者的历史战绩当论据:

"In some sense, scaling laws are like the most hated prediction of all time. Everybody always wants to say they're going to run out. They can't be like this. And yet it keeps going."
「从某种意义上说,扩展法则就像是历史上最被厌恶的预测。每个人总是想说它们会耗尽。它们不能就这样。然而,它们却一直在继续。」

他给创业者的操作建议是把"scaling laws are going to continue"直接内化成规划前提(为两年、四年后才可能/才经济的能力做设计)。但被问到"要继续 scaling unabated、最大的瓶颈是什么"时,他的答案把约束整个搬出了算法层:

"Transistors and then electrons in that order."
「晶体管,然后是电子,按这个顺序。」

——在他的框架里"模型触顶了吗"是个物理供给问题(芯片、能源、数据中心),不是算法回报问题。同时记录:他也埋了自己的撤退条款——设想两年后算力过剩的情形时,他主动说出"if we don't drive the cost curve down because we hit some sort of scaling wall, we could also get an oversupply"(如果因为撞上某种 scaling wall 而无法压低成本曲线,我们也可能过剩)——与 Casado 的"前提"条款同款:看多者自己列出了看空情形的触发器。

阵营 C · "pre-training 慢下来,但新的增量在 reasoning / post-training"——OpenAI / Poolside 立场

OpenAI 的 Isa Fulford / Christina Kim 把 GPT-5 的改进直接归功于 data quality + RL,而不是单纯参数 / 算力:

"If you compare it to O3's front-end coding capability, this is just totally next level. It feels very different. … The team just really cared about like nailing front-end. And that means like getting the best data, like thinking about the aesthetics of the model and all of these things."
「跟 O3 的前端编码能力比,这完全是另一个级别,感觉很不一样。……团队真的非常在意把前端做好。这意味着拿到最好的数据、认真想模型的审美这些事。」
"One thing that's interesting is with reinforcement learning, training a model to be good at a specific capability is very data efficient. You don't need that many examples to teach it something new."
「有意思的一点是,用强化学习把模型训得擅长某个具体能力,是非常 data efficient 的——你不需要很多例子就能教它一件新东西。」

Eiso Kant (Poolside 创始人)(2026-08 新增)是这一阵营第一个非 OpenAI 的声音,而且把"增量在 post-training"推进到一个更锋利的版本:增量甚至不在"智能",在"行为"。他们 118B(8B 激活)的 Laguna S 打过两倍大的模型,他转述应用研究联合负责人 Peng Ming 的内部观察:

"I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early and being way more persistent."
「我感觉 Laguna S 的很多提升并非来自更多的智能,而更多来自不同的行为——更多验证、更少想当然、不提前宣布胜利、并且执着得多。」
Eiso Kant(转述 Peng Ming)· Inside the Model Factory
"We are going to be able to squeeze so much more out of smaller models than I think we had imagined in the industry. Because yes, there's intelligence and larger models are more intelligent. No doubt about it. We should continue to scale up. But the behaviors of being really persistent, of being able to backtrack when you're wrong, of like understanding how to interact with your environment, show us that we can get a lot more out of it."
「我们能从小模型里榨出的东西,会比这个行业此前想象的多得多。因为——是的,智能存在,更大的模型更智能,这毫无疑问,我们也应该继续 scale up。但真正执着、错了能回溯、懂得与环境交互这些行为告诉我们:我们能从中拿到多得多的东西。」

注意他不是看空 scaling(他在同一集里说,若有无限算力"明天大概就有 AGI")——但这个发现让他第一次公开动摇了"更大总是更值"的部署逻辑,给出了本主题语料里第一条最优模型尺寸曲线的表述:

"So I know that at the very limit, I'm not going to use the world's largest model one day, quadrillion parameter, whatever crazy like skill we scale up to do a basic coding task. Already today, I'm starting to size down for certain tasks. So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work. At some point, we're at the peak. And after that, the return on investment of using a bigger model just doesn't make sense. Now, I think the question is, before I would have thought that peak was very far away."
「所以我知道,到极限处我也不会有一天用世界上最大的模型——千万亿参数、随便我们疯狂 scale 到什么规模——去做一个基础的编码任务。今天我已经开始为某些任务把模型往小里换。这意味着存在一个最优点。意味着知识工作随模型规模上升存在某条曲线:到某个点我们到达峰值,之后用更大模型的投资回报就不再合理。而我认为问题在于——以前我会以为那个峰值非常遥远。」

"以前我会以为那个峰值非常遥远"——这半句是真正的新信息:一个每天训模型的从业者在下调知识工作所需模型规模的估计。如果他对,frontier 尺寸竞赛的经济学(而非能力上限)会先于 scaling laws 本身触顶。

阵营 D · "Data 才是被严重低估的轴"——Datology / 数据派立场

Ari Morcos (Datology) 是少数把整套问题重新框架的人:

"Data is the most under-invested in area of research relative to its impact, and I don't think it's even close."
「相对于影响而言,data 是 ML 研究里投入最不足的一块——而且差距很大。」
"Even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data. Which is insane. We know that all data are not created equal, that garbaging garbage out is like the oldest adage in computer science."
「哪怕你去看 Kaplan、Chinchilla 以及所有其他这些 scaling laws 的工作,它们全都假设 IID(独立同分布)数据。这太离谱了。我们都知道'数据并非生而平等'——'垃圾进、垃圾出'是计算机科学里最老的格言。」
"Making the data better can be a massive compute multiplier. It can change the performance per dollar by orders of magnitude."
「把数据做得更好,可以成为巨大的算力倍增器——每美元性能可以变化好几个数量级。」

这条线把"plateau"问题转换成"我们一直在用错的资源花算力"。

阵营 E · "Capability ≠ Utility"——Brad Lightcap / Zelikman 的角度

Brad Lightcap 直接给出最强版本的"capability 已经远超 deployment"论点:

"You could stop progress right now. And I still think there's kind of a 10 or 20 year diffusion and innovation cycle that just to get it into the economy."
「就算现在停下进展,我还是认为有 10 到 20 年的扩散和创新周期——光是把它送进经济里就要这么久。」
"When you reduce the cost of something to zero, the demand for it goes up significantly."
「当你把某样东西的成本降到零,对它的需求会显著上升。」
Brad Lightcap · Uncapped #46

Eric Zelikman (humans&) 从另一边补了同样的判断:

"We have these incredibly smart models that are capable of so much, but they're not used for anywhere near what they're capable of."
「我们手上有这些极其聪明的模型,能干的事情很多——但它们的使用远远没到它们能做的水平。」

Zelikman 还给了一个不在其他阵营里、被忽视的瓶颈——情商(emotional intelligence)。他创办 humans& 的核心论点正是:当下模型常常"失败",瓶颈不在 IQ,而在缺乏情商和对人类价值的理解,于是"真正帮到人"的能力被卡住。(此条为其访谈要点转述。)

阵营 F · "先别问模型触没触顶,先问尺子还准不准"——Noam Brown 的测量派重构(2026-07 新增)

Noam Brown(OpenAI,推理范式的先驱之一) 把整场争论向后退了一步:他不直接回答"模型触顶了吗",而是论证用来判断触顶的仪器——benchmark 网格——已经失效。机制有两步。

第一步,能力不再是模型的静态属性,而是推理预算的函数:

"The problem is we're in a world now where the capability of the model is a function of how much money you put into it, basically. If you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. If you give it a budget of $10 million, it can do even more. And so at what budget should you evaluate these models? The policies that exist today don't really address that question."
「问题是我们现在的世界中,模型的能力基本上取决于你投入多少资金。如果你给它 10,000 美元的预算,它能做的比给它 10 美元时多得多。如果你给它 1000 万美元的预算,它能做的更多。那么,在什么预算下你应该评估这些模型?目前存在的政策并没有真正解决这个问题。」

第二步,行业发布模型时的"网格"不控制这个预算变量,于是纸面增量系统性缩水——这正是 5.5 发布时被质疑"没进步"的原因:

"The benchmark results are being presented in the wrong way. They're not controlling for the amount of test-time compute that is being used on that benchmark question. It turned out that 5.5 is just much more efficient with its thinking."
「基准以错误的方式呈现。他们没有控制在那个基准问题上使用的测试时计算量。结果表明,5.5 在思考上更加高效。」

这个机制直接打击语料里所有从"benchmark 只涨了几个百分点"推出 plateau 的判断(阵营 A 里 Vibe Check panel 的技术性观察就属于这一类)。值得注意的对照:Jack Morris 的"记忆容量触顶"是训练侧的信息论结果,不依赖推理预算——是阵营 A 里不受这条批评波及的那一条。

更进一步,当被要求"让模型跑到性能 plateau 再评估"时,Noam 的回答等于宣布"plateau 点"在实操上已不可达:

"The thing is, the point at which it plateaus is actually really far out these days. … 5.5 and other models … if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks. And so the point at which they plateau is simply too far out to reasonably test."
「问题是,平稳的点现在实际上非常遥远。……5.5 和其他模型如果合理地进行支撑,可以持续思考数周,甚至在某些基准上达到表现平稳之前。所以它们平稳的时刻实际上太远,不适合合理测试。」

他给了一个可查证的硬数字,以及一个更激进的推论——天花板在哪没人知道,因为没人跑够长

"The AISI in their evaluations has shown that the models continue to improve at 100 million tokens, if you run them for 100 million tokens, they're still improving at beyond that point."
「实际上 AISI 在其评估中显示这些模型在 1 亿个标记时仍在持续改进,如果你运行它们达到 1 亿个标记,它们在这一点之后仍在改进。」
"And so nobody actually knows what the ceiling of capabilities are for these models because nobody's actually run them for long enough to really tell."
「所以没人真正知道这些模型的能力上限是什么,因为没有人真的运行过足够长的时间来真正判断。」

那为什么行业不改测量方式?Noam 给的解释不是技术难度,而是博弈均衡:

"You kind of end up in this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out."
「所以你最终会陷入这种糟糕的平衡中,每个人都知道这是个糟糕的平衡,但没有人想要打破它。」

Mark Chen 从评估生产侧确认了同一件事——不是模型停了,是尺子饱和了:

"And we really are kind of in an evals crisis, right? Where all the really great evals that we all know, like growing up, like taking the SAT, those are all fully saturated. We really need to find good new ways to benchmark the models."
「我们确实正处于评估危机之中,对吧?所有我们都知道的,像小时候做 SAT 的那些优秀评估,都是完全饱和的。我们真的需要找到良好的新方法来基准测试模型。」
"There's this philosophy of once an eval is out in the world, then it's just already not a good eval."
「有一种观点是,一旦评估发布在世界上,它就已经不再是一个好的评估。」

Camp F 内部的限定同样锋利——这不是无限看多。Noam 自己承认有些能力轴对预算不敏感(事实检索类问题给一周也不会更好),research taste 也还不行:

"There are some benchmarks where the models will just not improve if they have more inference budget."
「有一些基准,如果模型有更多的推理预算,就不会改善。」
"One thing I see for research in particular is they don't have very good research taste right now."
「我认为目前在研究上,它们并没有很好地把握研究的品味。」

而且这套"能力=预算的函数"的世界观在他那里反过来变成反对"一夜智能爆炸"的论据——纸面低估进步与物理时间限速,是同一枚硬币的两面:

"If it requires so much test-time compute to unlock the full capabilities of the model, then that means you're bottlenecked by time. Things can only go so fast because the models need to run for long enough to actually do something really, really powerful. Time itself becomes a bottleneck to what we can do."
「如果解锁模型的全部能力需要如此多的测试计算,那么这意味着你受到时间的限制。事情的速度只能如此快,因为模型需要运行足够长的时间才能真正做出非常强大的事情。时间本身成为我们能做的事情的一个瓶颈。」

这一阵营留下了语料里最可核销的一批预测和数字。Noam 的扑克 solver(他的博士论文课题):

"And I wouldn't be surprised if, you know, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go."
「如果六个月或一年后,模型能够零样本完成整个扑克解算器,我不会感到惊讶,基本上是我整个博士论文一次性完成。」

Erdős 单位距离猜想的反证成本(同一任务的价格随代际塌缩——这是"能力=预算函数"最直观的量化):

"The cost of disproving the Erdos unit distance conjecture drops by like 10 or 100x with every model release cycle, probably in some cases more."
「所以反驳 Erdős 单位距离猜想的成本在每次模型发布周期中下降了大约 10 或 100 倍,在某些情况下可能更多。」

Mark Chen 一侧对应的时间表——三年路线图的终点:

"When we look at our kind of three-year roadmap, the end goal that we want to reach is one where the models are just doing end-to-end research."
「当我们看我们的三年路线图时,我们想要达到的最终目标是模型能够进行端到端的研究。」

阵营 G · "触顶的不是 scaling 假说,是 Transformer 架构本身"——Core Automation 的架构派(2026-08 新增)

Jerry Tworek(前 OpenAI 副总裁,Strawberry/推理团队负责人)与 Rohan Anil(前 Gemini 预训练负责人之一,Google Brain / Anthropic) 开了一个所有现存阵营都没占的位置:他们接受 scaling 成功了、也接受它到头了——但触顶的对象既不是算力回报也不是数据,而是 Transformer 这个架构能表达的东西。Jerry 的开局陈述:

"We got really, really good at training really, really big models. We mastered two algorithms. We mastered pre-training at a large scale and we mastered reinforcement learning at a large scale. And I'm asking myself a lot, what is next in machine learning? And I think at this moment, what the bottleneck is to better models and to smarter systems is the architecture itself."
「我们在训练非常非常大的模型方面变得非常出色。我们掌握了两种算法。我们在大规模预训练和大规模强化学习方面都取得了突破。我经常问自己,机器学习的下一步是什么?我认为此时模型和智能系统的瓶颈在于架构本身。」

支撑这个立场的是语料里最贵的一段个人证词——他是 OpenAI 推理范式的当事人,讲了自己作为"RL maximalist"的幻灭:

"I was thinking, here we are, if you ask Jerry in 2024, when do we get AGI? I would say 2025 will be that year. This is where we solve everything. And I saw us training model after model. This model was getting better and better. All the benchmarks scores were going up. And did we also solve all the real world tasks at that moment? Unfortunately, unfortunately not."
「我在想,如果你在 2024 年问杰瑞,我们什么时候能实现 AGI?我会说 2025 年将是那一年。这是我们解决所有问题的时候。我看到我们在训练一个又一个模型。这个模型变得越来越好。所有的基准分数都在上升。那时我们是否解决了所有现实世界中的任务?不幸的是,不是的。」

他对"为什么分数涨了、现实没解决"给出的机制,是对整个 benchmark 体系的另一面指控——与 Noam Brown 方向相反

"I realized there was this bit of distinction as all the benchmarks that we are evaluating our models, they were essentially the same thing as we were training the models on. Like all the evals and training tasks are the same size of the coin, but the real world distribution and real world task is much messier, much murkier, much more different."
「我意识到,所有评估我们模型的基准与我们训练模型的任务基本上是相同的。所有的评估和训练任务都是同一枚硬币的两面,但现实世界分布和实际任务是非常混乱、模糊,非常不同的。」

把这段与阵营 F 并排看:Noam Brown 说"网格不控推理预算,基准低估了能力";Jerry Tworek 说"评测与训练同源,基准高估了现实迁移"。两人都是 OpenAI 推理路线的核心当事人,两个失真方向同时为真是可能的——但那样一来,benchmark 作为"触顶了吗"的证据就双向失效了(详见"都没说透的"新条目)。

Jerry 由此推出的缺口是持续学习:in-context learning 太小(他用 Codex 约 20 分钟就得 compact),持续微调有灾难性遗忘——"the models are being trained in the lab. And are being deployed in the real world. That is the fundamental tension that is there."(模型在实验室里训练、在现实世界里部署,这是根本张力所在)。他的思想实验把"静态模型"的天花板讲得最直观:

"And then if we ever stop training that model, what would happen? The question we're asking often and thinking about transformer, what would happen if OpenAI and Antropic stopped training new models and we got the transformer we have today and say, this is it. This is the best model we have. Months pass, years pass and the model is getting less and less useful."
「然后如果我们停止训练那个模型,会发生什么?我们常常在思考变压器的问题,如果 OpenAI 和 Anthropic 停止训练新模型,我们得到了今天的变压器并说:。这就是它了。这是我们拥有的最好的模型。几个月过去,几年过去,模型变得越来越无用。」

Rohan Anil 补的是技术侧的两根梁。其一,深度:

"Most transformers that we train are quite shallow. It's at most 100 layers deep. It's called deep learning because you wanted deeper representations. There has been experiments on going into depth, but no one has actually shown us learning extremely deep representations."
「我们训练的大多数变压器相当浅显。深度最多为 100 层。它被称为深度学习,因为你想要更深的表示。在深入研究方面已经有一些实验,但没有人实质性地向我们展示过如何学习极其深的表示。」

——chain-of-thought 与推理时 scaling 在他看来是给浅架构续命的"band-aid"(逐 token 自回归地借序列长度换计算深度)。其二,也是对本主题第一条共识最直接的回应——他接受对数线性的数学,但宣布这把尺子本身问错了问题

"And every time we increase compute in log scale, we get this epsilon more improvement in these metrics. This is, I think, this is fine for building the prior. I think this is the wrong way to look at the problem. We should be looking at the end-to-end. What are we training these models for?"
「每次我们以对数比例增加计算时,在这些指标上就会得到更多的 epsilon 改进。我认为,这对于构建先前模型是可以的。我认为这种看问题的方式是错误的。我们应该关注端到端。我们训练这些模型的目的是什么?」

对照 Bob McGrew 的"收益递减是固有的":Rohan 不否认那条曲线,他否认 perplexity 是值得画曲线的 y 轴。这是把"plateau 是尺子的问题"(阵营 F)推广到了训练侧指标。

值得记录的跨阵营呼应:Sam Altman 在同月的访谈里,列举当前模型"还缺什么"时点的正是同一块——

"Even some of the real skeptics have said to me in recent days or recent weeks, I guess, I think GPT 5.6 has been out for me two weeks. They're like, okay, It's very hard for me to say what I want from this model that it can't do. … The model also, although brilliant, is still not learning continuously as it goes."
「甚至一些真正怀疑的人最近告诉我,或者最近几周,我想 GPT 5.6 已经发布两周了。他们说,好吧,我很难说出我想从这个模型中获得什么,它无法做到。……虽然这个模型很聪明,但仍然没有持续学习。」

在位者与出走者对"缺持续学习"达成一致,分歧只在这算不算"触顶":Altman 视之为待办清单上的一项,Jerry/Rohan 视之为现架构原则上做不到的事。利益结构同样对称计入:Core Automation 的融资叙事依赖"Transformer 到头了"成立,正如 OpenAI 的估值依赖它不成立——他们指出的经济锁定(大实验室因 Transformer 可盈利、竞争白热化而无心投资替代架构)本身可信,但方向盘握在谁手里,谁就有理由把路况说得更极端。

都没说透的

"When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the entire web. But for RL, you have, you know, X millions of tasks that you are going to be training on. And so you cannot blow up your batch size massively, which means that you actually can't scale compute to a certain extent with RL the same way you could scale compute with like pre-training."
「预训练时你有整个互联网,batch size 可以很灵活地放大。但 RL 只有几百万量级的任务可训练,所以你没法大幅扩 batch size——这意味着在某种程度上,你无法像预训练那样用堆算力的方式去 scale RL。」

他明说"My biggest wall-clock bottleneck right now is RL time"(我现在最大的墙钟瓶颈就是 RL 时间)。也就是说:RL 轴不是被回报递减卡住,是被任务供给和 batch size 卡住——砸算力这条老路在这条轴上物理性失效。加上 Jerry Tworek 的当事人证词(RL 规模化后基准全涨、现实任务未解),阵营 C 的"新轴"从两个方向被收窄。三位从业者(Joelle、Eiso、Jerry)现在指向同一个结论的不同侧面,这条曾经的乐观轴该重新定价了。

我的看法

判断(不是事实):上一版的核心判断保留——pre-training 单轴回报变缓的共识依然最强(这一版 Babushkin 的正面确认让它更硬:终于有亲手跑过 frontier 预训练的人把话说满,而 Altman 的反对与 Mark Chen 一样,靠的是历史归纳而非机制),新增量在 post-training/RL、data quality、deployment 三条轴上分裂,"触顶了吗"被错误框定为单一维度。"plateau 感里有相当份额是测量假象"这条判断保留,但这一版必须双向化:Jerry Tworek 的"评测与训练同一枚硬币"表明失真不只有低估一个方向——benchmark 可能同时低估极限能力(不控预算)和高估现实迁移(与训练同分布)。这不推翻"以现行公开测量方式,无法判定"的结论,反而把它锁得更死:尺子不只刻度不准,正反两面都被人指为弯的。综合结论维持:不是"没触顶",是"测不出来"——且 2026-08 之后我会加一句:"触顶"最可能的形态已经不是能力停涨,而是(a)架构性缺口(持续学习)让能力增长与现实价值脱钩,或(b)最优尺寸经济学让前沿能力与买方需求脱钩。这两种"软触顶"都不出现在传统 plateau 叙事里,但新料里最有分量的三个声音(Jerry/Rohan、Eiso、Babushkin)全都指向它们。利益折价这一版必须对称化:旧版折价 OpenAI 在位者,这版同样折价新架构/分布式创业者——语料里已经没有无持仓的证人,我的权重只能给机制的可检验性,不能给资历。应用价值增长持续至少 5–10 年的判断不变,且又加两票:Altman 的"回到 2019 展示今天的模型,所有人都会说经济将被颠覆——而这并没有发生"是在位者对 capability≠utility 的自供,Eiso 的"行为增量"证明部署价值还有非尺寸来源可挖。

把握程度:应用价值持续增长——中等偏高(不变)。"plateau 感有相当份额是测量假象(现为双向失真版)"——中等(不变,理由换了:机制更完整,但两个方向的证词都来自持仓者)。"pre-training 单轴回报变缓"——从接近事实下调半档为强共识但已有实名两边对峙。新增:"RL 训练轴受任务供给/batch size 约束、无法靠堆算力延续"——中等偏高:Eiso 给出的是工程约束不是观点,Joelle/Jerry 从不同侧面独立指向同一处,且没有任何新访谈正面反驳它。可证伪点在原有两条(按预算控制的第三方对比;实验室按 Noam 方式出图)之上新增一条:若 24 个月内"不经实验室回炉的持续学习"在非 Transformer 架构上率先做到生产可用,阵营 G 的"架构触顶"论升权为主线;若它在 Transformer 的增量修补内被解决、或 Transformer + 更大 RL 预算继续产出代际跃迁且现实迁移同步改善,则降权。

还想知道什么

取材