主题综述

AI 能不能真正发现新科学 · Can AI Actually Discover New Science

主题综述

更新日志

主流共识

九篇访谈横跨基础模型 CEO(Altman)、AI 分析者(Labenz)、投资人(Deedy Das)、前沿实验室推理研究员(OpenAI 的 Wei/Wu/Chen),以及三家把"让 AI 做科学"当公司主业的团队(Radical AI 的 Krause、OpenAI for Science 的 Weil、Lila Sciences)。在一个天然容易两极化的话题上,他们在五件事上高度一致。

第一条共识:AI 已从"加速科学家"跨向"偶发拓展前沿",而且拿得出具体战果。 Kevin Weil 把整个命题压成一句口号——把 2050 年的科学提前到 2030 年:

"The models can now solve problems that humans have never solved before, going beyond the frontier of human knowledge. That's how AI, I think, and AGI will really change our lives. Why not try and accelerate science, bring about the science of 2050, but in 2030 instead?"
「这些模型现在能够解决人类从未解决过的问题,超越人类知识的边界。这就是我认为人工智能和通用人工智能将真正改变我们生活的方式。为什么不尝试加速科学发展,提前到 2030 年实现 2050 年的科学呢?」

他把当下定位在能力跃迁的"中间阶段"——不能做→勉强做→擅长做,这个模式正在前沿科学重演:

"you go very quickly from models could never do this thing … Models can just barely do this thing, and it kind of sucks at it, and it's wrong most of the time … and then six to 12 months later, it's like, models are great at this thing … We are clearly in that middle phase with frontier science and AI where you have all these glimmers of like, wow, it can do something that we never thought AI could do. So what more interesting place to apply it than science?"
「你很快就会从 "模型永远无法做到这件事" 变成……模型几乎可以做到这件事,虽然它很糟糕,而且大部分时间都是错的……六到十二个月后,情况就变成了,模型在这件事上表现很好……我们显然处于前沿科学和人工智能的中间阶段,那里有很多闪光点……那么还有什么比科学更有趣的应用场景呢?」

Nathan Labenz 把 GPT-4 和新一代之间的落差说成"天壤之别"——多家公司的纯推理模型无工具就拿下了 IMO 金牌:

"A big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-4 could do with math, right?"
「过去几周的一个重要进展是,一些公司推出了纯推理模型,无需任何工具即可获得IMO金牌。而且,你知道,这与GPT-4在数学方面的表现相比,简直是天壤之别,对吧?」

他强调这是质变而非量变——这些结果"以前根本没人知道",GPT-4 无法真正推动人类知识的前沿:

"And these are things that like literally nobody knew before. And GPT-4 just wasn't doing that. You know, I mean, these are qualitatively new capabilities… I would say, in summary, GPT-4 was not able to push the actual frontier of human knowledge."
「这些是以前根本没人知道的东西。GPT-4根本做不到这一点。你知道,这些是质的新能力……总而言之,GPT-4无法真正推动人类知识的前沿。」

OpenAI 推理团队交出的是数学史上的硬战果:一个 80 年的 Erdős 单位距离猜想被通用推理模型反驳——注意这是上世纪 Erdős 悬赏 500 美元的问题:

"So our models last week, they were able to produce a disproof of the unit distance conjecture due to ERDOS. And this was an 80-year-old open problem in the field of combinatorial geometry…"
「所以,我们上周的模型能够给出对 ERDOS 提出的单位距离猜想的反驳。而这是一个在组合几何领域存在了 80 年的未解问题……」

而且这不是复述已知——模型推翻了长期被当成最优解的方格构造:

"What we found was essentially that the optimal solution to having as many distance one points on the plane was to arrange them in a square grid. And what the model proved was that the square grid was not actually close to optimal at all, and that you can do much better with a different construction using a lot of high-powered number theory."
「我们发现基本上是,平面上尽可能多地放置距离为一的点的最佳解决方案是将它们排列成一个方格。而模型证明方格实际上并不接近最佳,而是可以通过使用大量强大的数论做得更好。」

Sam Altman 从模型迭代给出同一判断——从 GPT-5.5 起,顶尖科学家开始报告能靠模型想出更好的点子,模型已能做出"虽小但重要的发现":

"Starting with the models of a few months ago, but really now with 5.5, the models have gotten smart enough that excellent scientists are saying, I am able to Figure out better ideas."
「从几个月前的模型开始,特别是现在的 5.5 版本,模型已经足够聪明,以至于优秀的科学家们都在说,我能够想出更好的点子。」
"The models are able to make some small but important discoveries. And the pace of science is going to increase. Eventually, we'll have, you know, automated labs and robots and, you know, building who knows what. And we'll be able to do science much faster. But if we can start doing like, a decade of science, of what it would have taken us in the old world in a year, the compounding effect there and what we'll be able to do and discover will just be extremely great."
「模型能够做出一些虽小但重要的发现。科学的进步速度将会加快。最终,我们将拥有自动化实验室和机器人……如果我们能用一年时间完成过去需要十年才能完成的科学研究,这种复合效应所带来的成果和发现将极其巨大。」

第二条共识:但现阶段主流仍是 copilot / 加速器,"自主发现"是下一个里程碑而非现状。 Altman 划得最清——你不能对 ChatGPT 说"找出新物理"并指望它奏效,当前是 copilot 式,但已听到生物学家的轶事:

"Yeah, there's definitely not. Like you definitely can't go say like, hey, ChatGPT, figure out new physics and expect that to work. So I think it is currently Copilot-like, but I've heard like anecdotal reports from biologists where it's like, wow, it really did figure out an idea. I had to develop it a little bit more, but it made like a fundamental leap."
「是的,绝对不是。比如你肯定不能说,嘿,ChatGPT,找出新的物理学,并期望它能奏效。所以我认为它目前类似于 Copilot,但我听到了一些生物学家的传闻,说,哇,它真的想出了一个主意。我不得不进一步发展它,但它做出了一个根本性的飞跃。」

即便还没有自主科学,一个科学家用 O3 效率提高三倍已经很重要——这是"加速"而非"发现":

"Both. I mean, you already hear scientists who say they're faster with AI. Like we don't have AI maybe autonomously doing science. But if a human scientist is three times as productive using O3, that's still a pretty big deal. And then as that keeps going and the AI can like autonomously do some science, figure out novel physics."
「两者都有。我的意思是,你已经听到科学家说他们使用人工智能后速度更快了。比如我们还没有人工智能自主地进行科学研究。但是如果一个人类科学家使用 O3 的效率提高了三倍,那仍然是一件非常重要的事情。然后随着这种情况的持续发展,人工智能可以自主地进行一些科学研究,找出新的物理学。」

Labenz 给了最诚实的限定——那个著名的 AI co-scientist 案例里,模型对困扰病毒学界多年的开放问题给出了假设,恰好与科学家未发表的实验答案完全吻合,但它自己无法验证:

"In one particularly famous kind of notorious case, it came up with a Hypothesis, which it wasn't able to verify because it doesn't have direct access to actually run the experiments in the lab, but it came up with a hypothesis to some open problem in virology that had stumped scientists for years and it just so happened that they had also recently figured out the answer but not yet published their results. And so there was this confluence where the scientists had experimentally verified and Gemini, in the form of this AI co-scientist, came up with exactly the right answer."
「它提出了一个假设,但它自己无法验证,因为它没有能力直接在实验室里运行实验……但它对病毒学中一个困扰科学家多年的开放问题提出了假设,而恰好科学家们最近也得出了答案但尚未发表,于是出现了这样的巧合:科学家已经通过实验验证了答案,而Gemini这个AI co-scientist给出了完全正确的答案。」

而真正全新的发现,他坦承目前仍然稀少,只是"开始偶发"——这本身就是大事:

"To my knowledge, I don't know that it ever discovered anything new. It's still not easy to get that kind of output from a GPT-5 or a Gemini 2.5 or a Claude Opus 4 or whatever, but it's starting to happen sometimes. And that in and of itself is a huge deal."
「据我所知,我不知道它是否曾经发现过任何全新的东西。想从GPT-5、Gemini 2.5或Claude Opus 4这样的模型中获得那种产出仍然不容易,但它已经开始偶尔发生了。这本身就是一件大事。」

Lila 的 Andy Beam 也承认距离"完全开放式自由实验"还有距离——尽管模型在基因编辑协议设计上零样本已达约 80%(人类零样本为 0%):

"The model gets like 80% of that zero shot. Humans get 0% of that zero shot. Are we doing like fully open-ended free-form experimentation now? I mean, no, not yet. That is the goal. But we're building towards that. That is the end state that we want to be in."
「模型在零样本情况下能够达到大约 80% 的效果。人类在零样本情况下的效果为 0%。我们现在是不是在进行完全开放式的自由实验?我的意思是,不,还没有。这是我们的目标。但我们正在朝这个方向努力。」

第三条共识:终局形态是闭环——推理 + 机器人实验室 / self-driving lab / 24-7 RL 循环。 Weil 把架构画得最细——紧循环(模型-仿真)套在长循环(现实世界机器人实验室)里,可横向扩展:

"the science of the future will definitely involve robotic labs and reinforcement learning loops that go through the real world where the model is thinking, maybe running a simulation, thinking some more, refining the experiment … and then sending that to a bunch of robotic labs, which, by the way, you can scale horizontally, having the experiments run in real life, the results come back to the model, the model thinks, runs more simulations, thinks, and you have this, you have like multiple loops."
「未来的科学一定会涉及机器人实验室和强化学习循环,这些循环会在现实世界中进行,模型在思考,可能运行一个仿真,继续思考,完善它可以使用最佳参数进行的实验,然后将这些发送到一系列机器人实验室……这些实验室可以进行横向扩展,在现实生活中运行实验,结果返回给模型,模型思考,运行更多仿真,思考……你会有多个循环。」

关键是这些机器人实验室 24/7 运行,绕开了人类研究者需要睡觉、休息的生理限制:

"you have robotic labs that you can scale horizontally that can run 24 hours a day. They're not grad students pipetting things that need to take breaks and sleep. And then the grad students can do things that are much more leveraging of what makes us human than pipetting things. I'm quite optimistic about where this goes … to accelerate the pace of science very meaningfully."
「你有可以横向扩展的机器人实验室,可以全天候运行。他们不是需要休息和睡觉的研究生去移液。然后研究生们可以做一些更能发挥人性特质的事情……我对这一切的发展方向……非常乐观……非常有意义地加速科学的发展步伐。」

Lila 把这个闭环重新定义成"科学的规模化验证器"——用自然和实验当可验证的奖励信号,在大规模后训练里推动推理模型的边界:

"At Lila, what we believe is that actually science, running the scientific method and using nature and experiments as verifier is like the ultimate version of that. ... what we're building, we'll talk about these things that we call AI science factories. They are scaled verifiers for science so that we can do post-training at scale and push out the frontier of what reasoning models are capable of."
「在 Lila,我们相信,实际上,科学、运行科学方法以及使用自然和实验作为验证者就像是这一切的终极版本。……我们正在构建的东西,就是我们称之为 AI 科学工厂的东西。它们是科学的规模化验证器,以便我们能够进行大规模的后训练,并推动推理模型的能力边界。」

Altman 的版本是一个思想实验——花一千亿美元建粒子加速器,让 AI 看数据、决定做什么实验:

"I wonder about this. If you could like build AI, a hundred billion dollar particle accelerator. … And say you'd make decisions. You look at the data. You tell us like, you know, what experiments to run and we'll go like find the stuff and do it. Or you spend a hundred billion dollars like building an infrastructure to like connect into the economy. Which of those will it have an easier time doing something remarkable with?"
「我想知道这个。如果你可以用人工智能建造一个价值一千亿美元的粒子加速器。并且说你会做决定。你看看数据。你告诉我们要做什么实验,我们会去找材料并完成它。或者你花费一千亿美金来构建连接到经济的基础设施。哪一个更容易做出非凡的成就?」

第四条共识:"实验才是护城河,模型不是"。 这句话几乎是逐字的共识。Radical AI 的 Krause 说得最直白——五年内模型都会开源,实验才是壁垒:

"We actually think in five years most models will be open source. … we think experiments are the moat, not the model itself."
「我们实际上认为在五年内,大多数模型将会是开源的。……我们认为实验是护城河,而不仅仅是模型本身。」

Lila 用"代币生成器"给同一句话换了个说法——实验平台而非模型本身,才是可持续优势的来源:

"The lab platform is the token generator... that ultimately is the moat for Lila is that once that continues to scale, the amount of data that we can generate, both per unit time, but per unit square foot will go up. And that feeds back into the model to make it smarter, that then suggests the next experiment to do."
「实验室平台是代币生成器……最终构成了丽拉的护城河,一旦它继续扩展,我们可以生成的数据量,无论是每单位时间,还是每平方英尺都会增加。这会回馈到模型中,使其更智能,随后建议进行下一个实验。」

Krause 把这个护城河的物理根据说透——材料的"地面真相"就是材料本身,必须造出来、测出来,绕不过:

"in materials, the ground truth is the material itself. You have to be able to make it. You have to be able to test it and characterize it."
「在材料方面,事实真相就是材料本身。你必须能够制造它。你必须能够测试并表征它。」

第五条共识:领域进度不均,分野在 experiment-constrained vs compute-constrained、有没有统一表征。 Weil 解释为什么数学、物理、理论 CS 先行——可以完全在计算机里闭环优化:

"We started thinking about math, physics, theoretical computer science, because you can do everything in Silico. You have closed-loop systems that you can optimize."
「我们开始考虑数学、物理、理论计算机科学,因为你可以在 Silico 中做任何事情。你可以优化的闭环系统。」

Altman 举了另一个例子——天体物理可能最先出现自主发现,因为数据海量而人类博士稀缺,这是数据受限而非算力受限的典型:

"And I think the physics is a cleaner problem. You know, I think if you could get like new high-energy physics data and then AI the ability to like run experiments, I think that's like a cleaner problem. I've heard people say that... they expect the first area of science where AI makes autonomous new discoveries to be astrophysics because there's just mountains of data and we don't have enough PhDs to look at it. And maybe it's not that hard to figure out new stuff, but I don't really know."
「我认为物理学是一个更清晰的问题。我认为,如果你能获得新的高能物理数据,然后 AI 能够运行实验,我认为这是一个更清晰的问题。我听人说他们期待科学的第一个领域...他们期望 AI 做出自主新发现的第一个科学领域是天体物理学,因为那里有大量的数据,而我们没有足够的博士来研究它。也许弄清楚新的东西并不难,但我真的不知道。」

Krause 把材料科学摆到光谱的另一端——它是实验受限,不是算力受限:

"we don't, we're not compute constrained in the materials industry... Yes, we're experiment constrained."
「我们在材料行业并不受计算能力的限制……是的,我们的实验受到限制。」

主持人直接抛出"材料领域没有 AlphaFold 时刻"的命题,Krause 同意——子问题(如显微图像)可以有局部的 AlphaFold 时刻,但端到端从假设到成品做不到:

"Okay, I do. I think that I think you can have alpha fold moments for specific areas of materials like microscopy... What I don't think you can do is go from like, I have this new hypothesis to, oh my gosh, I have a new material. It's scaled. It's done. It's in products in your iPhone. You can't do that. I would agree with that statement."
「好吧,我同意。我认为你可以在特定材料领域有 Alpha Fold 的时刻,比如显微镜领域……我认为你做不到的是,从我有这个新假设,到哦我的天,我有了新材料。这已经扩大了。已经完成了。它在你手机里的产品中。你无法做到这一点。我同意这个说法。」

Lila 的 Andy Beam 给出机制解释——材料缺乏像生物学"中心法则"那样的统一原则,实验室测试对真实寿命只是部分可预测:

"For me, materials as a subject is interesting because there's not a unifying principle like the central dogma. Like materials means lots of different things... the testing that you do in the lab is only partially predictive of like the lifetime of how that material will be used. And the math is harder."
「对我来说,材料作为一个学科很有趣,因为没有像中央教条这样统一的原则。像材料代表着很多不同的事物……实验室中进行的测试只是部分地预测材料如何使用的寿命。而且数学更难。」

Rafa Gómez-Bombarelli 补上更硬的技术原因——物理仿真"预测性不够"(sim-to-real gap),所以材料只能靠 self-driving lab 做物理验证,不能纯计算预测:

"For us in physics-based simulations is that, you know, we do molecular simulations of gooey stuff, we do electronic structure simulations of hard stuff, and they're okay, but they're not predictive enough. So this is the reason why... maybe we would have had to make a self-driving lab for materials because we would have been able to just predict."
「对于我们在基于物理的仿真中来说……我们进行黏稠物质的分子仿真,进行坚硬物质的电子结构仿真,它们还可以,但预测性不足。所以这就是原因……若不是因为这个,也许我们会不得不为材料创建一个自动驾驶实验室,因为我们本可以直接预测。」

分歧在哪

共识到此为止。一旦追问"这到底算不算真发现、纯推理能走多远、需要什么前提",九个人就裂开了。

一、AI 今天做的,是"真新科学"还是"高级 remix"?

这是最锋利的一处对撞,而且是本主题的题眼。一端是 Altman 的极大值——他几乎把"能自主发现新科学"当成超级智能的定义:

"If we had a system that was Capable of either doing autonomous discovery of new science or greatly increasing the capability of people using the tool to discover new science. That would feel like kind of almost definitionally super intelligence to me and be a wonderful thing for the world, I think."
「如果我们有一个系统能够自主发现新的科学,或者大大提高人们使用该工具发现新科学的能力。对我来说,这几乎可以被定义为超级智能,而且我认为这对世界来说是一件美好的事情。」

另一端是 Deedy Das 的极小值——今天 AI 的"新发现"大多是"伪新颖",人类花足够时间也能找到;那是知识树上的一根小树枝,不是一根新枝干:

"So can we actually do real, and there are some novel discoveries that AIs can do today, but they're kind of pseudo-novel in the sense that like, yeah, if you really put a human to go discover that and look at the data for that long, they'd probably discover it. It's just not that interesting type of discovery. And the way to think about it is like, if the branch of knowledge is like this, it's like kind of a twig on one side. It's cool, but it's not like a branch."
「那么我们真的能做到吗?今天人工智能可以做一些新的发现,但它们在某种意义上是伪新颖的,因为,是的,如果你真的让人类去发现它,并长时间地查看数据,他们可能也会发现它。这只是一个不太有趣的发现类型。考虑它的方式是,如果知识的分支像这样,它就像一侧的一根小树枝。这很酷,但它不像一个分支。」

她把真正的门槛立在"发明知识的分支",并明确说当前技术不够、未来不可知:

"How do we get to like, hey, can we invent branches of knowledge? So some people think we can never get there. I know the current techniques are not good enough to get there. So my TLDR to that is I do not know the future. I do not know how to answer things like AGI 2027."
「我们如何才能创造知识的分支?所以有些人认为我们永远无法到达那里。我知道目前的技术还不够好,无法到达那里。所以我对这件事的总结是我不知道未来会怎样。我不知道如何回答像 AGI 2027 这样的问题。」

最耐人寻味的是:亲手攻下 Erdős 的 OpenAI 团队,反而站在离 Deedy 更近的位置。Lijie Chen 明确划界——AI 目前无法为数学构建新理论,它做的是从不同领域抓取想法、赋能人类:

"Currently, it seems AI cannot build a new theory for math. For example, but I guess humans, once they have the help of AI, they can just grab all the ideas from distinct fields of math. I think they can empower humans way more."
「目前看来,人工智能似乎无法为数学创造新的理论。例如,我想人类一旦获得人工智能的帮助,他们可以从不同的数学领域获取所有的想法。我认为他们可以更大程度地赋能人类。」

他把核心张力点得最准——组合已有想法(哪怕方式非常新颖复杂)和从零生成全新想法,是两回事,后者还没见过:

"…It currently seems AI is trying to combine ideas from different fields and of course in a very novel and sophisticated way. But can AI actually generate completely new ideas from scratch? I mean, that's something we haven't really seen concretely in AI. And that's something I maybe want to see next happening."
「目前似乎 AI 正在尝试结合来自不同领域的想法,当然是以一种非常新颖和复杂的方式。但是 AI 是否真的能从零开始生成全新的想法?我是说,这是我们在 AI 中还没有看到的具体事情。我想我可能想看到接下来发生的事情。」

他甚至把那份 125 页思维链拆开看——有创造性痕迹,但最终解法本质仍是"把所有东西结合起来":

"I think so. You know, even in this Erdős problem, I mean, I think if you look at the chain of thought, which is like 125 pages, I think some of the thoughts are pretty creative, although they didn't work out. I mean, the final idea is more like combining all the stuff, but it has some creative thoughts."
「我想是的。你知道,即便在这个厄尔多斯问题中,我的意思是,如果你看看思路的链条,长达 125 页,我认为其中一些思路相当有创造性,尽管它们没有结果。我的意思是,最终的想法更像是把所有的东西结合在一起,但它有一些创造性的想法。」

Alexander Wei 把这层张力说得最精确——把类域理论嫁接到组合几何"以前没人真正做过",需要洞察和创造力(这是真的 remix,但是极难的 remix);可"发明新的数学方法"是数年到数十年的过程,不同于解题能力每几个月翻倍的摩尔定律:

"So for some context, like the proof is like well above my own mathematical pay grade. But like just at a high level, my understanding was that… this idea of taking class field theory and applying it to problems in combinatorial geometry hadn't really been done before, though some people knew that there could be this bridge between these two fields. Being able to do that and execute it requires… quite a bit of insight and creativity."
「作为一些背景,这个证明的水平远超我自己的数学能力。但从高层次来看,我的理解是,这种将类域理论应用于组合几何问题的想法在之前并没有真正被做过……建立联系需要较大的洞察力和创造力。」
"How I would think about it is, we see this like… Moore's law for the time horizon at which these models are effective… But I think for inventing like new ways of doing mathematics that's much more like a years or decades long process. And so I think… it'll still take a bit of time for that exponential to get there."
「我会这样考虑,我们看到这种情况,就像……摩尔定律关于这些模型有效的时间范围……但是我认为发明新的数学方法的过程则更像是需要几年或几十年的过程。所以我认为,这个指数增长可能还需要一点时间才能到达。」

但"remix"不等于"没拓展知识"——同一个团队给出了反证:数学家们一周内就用同样的思路反驳了另一个关于实数的乘积猜想,说明这不是孤立巧合,而是能外溢的真实知识:

"Yes, some of the mathematicians that we asked to review the proof, together with collaborators, they actually used the idea to disprove some product conjecture, but for real numbers. I think that's like one very good example. AI can crack down on important questions and give us ideas that we can apply elsewhere."
「是的,我们请来的几位数学家与合作者一起复审证明,他们实际上用了这个想法来反驳某些乘积猜想,但针对实数。我认为这就是一个很好的例子。人工智能能够解答重要问题并提供其他地方可以应用的洞见。」

Labenz 站在"已越过前沿"这一端最坚决——他给的不是数学而是生物学的硬证据:MIT 用狭义专用模型造出了机制全新、对耐药菌有效的抗生素:

"We're at the point now with these biology models and material science models where they're kind of like the image generation models of a couple years ago… it's been enough for this group at MIT to use some of these relatively… narrow purpose-built biology models and create totally new antibiotics. New in the sense that they have a new mechanism of action… Notably, they do work on antibiotic-resistant bacteria."
「我们现在处于这样一个阶段:这些生物学模型和材料科学模型,有点像几年前的图像生成模型……即便如此,这已经足以让MIT的这个小组使用一些相对狭义的专用生物学模型,创造出全新的抗生素。新,是指它们有一种新的作用机制……值得注意的是,它们对耐抗生素细菌确实有效。」

他的第二个证据是斯坦福 Virtual Lab——多个专家型 agent 相互辩论、调用 AlphaFold 类工具模拟分子互作,产出了针对新冠新毒株的新疗法:

"And basically this was an AI agent that could spin up other AI agents… you have agents using the AlphaFold type… to say, okay, well, can we simulate how this would interact with that? Agents are running that loop, and they were able to get this language model agent with specialized tool system to generate new treatments for novel strains of COVID that had kind of escaped the previous treatments."
「这基本上是一个可以启动其他AI代理的AI代理……他们使用AlphaFold这类工具来模拟这个东西会如何与那个东西相互作用。代理们运行着这个循环,他们得以让这个具备专业工具系统的语言模型代理,为已经逃脱先前疗法的新冠新毒株生成新的疗法。」

Krause 也被主持人直接追问过这道题——你是在已知空间里挑新颖排列,还是真在推前沿?他选后者,并给出实证规模:过去五六个月造了 1200 种合金,其中 300 种是文献中从未出现过的全新家族:

"So that brings up like a follow-on question about how much are you just sort of optimizing within a well-understood space and you're picking permutations that are novel versus trying things, experiments that really are pushing the frontier of science?"
「这引出了一个关于你在已知空间内优化的程度的问题,你是在选择新颖的不同组合,还是在尝试新的实验,真正推动科学的前沿?」
"The latter, we really are making new materials that push the frontier."
「后者,我们确实在制造推动前沿的新材料。」

Weil 站在中间,也最克制——2026 年 1 月一个月里十来个开放数学难题被解出(多数由 GPT-5.2、少数由 Gemini),模型确实越过了人类知识边界,但他"不敢说"它在解人类原则上解不了的问题,只是人类还没解、模型先做到:

"we've now seen there been I don't know what 10 or 12 just in January, 10 or 12 open mathematics problems solved, mostly by GPT 5.2, now a few recently by Gemini. Models are going beyond the frontier of human knowledge. And I wouldn't claim yet that they are solving problems that humans can't. I think if you took enough people and, you know, applied enough mathematicians towards some of these problems, they would have figured it out. But they had not figured it out yet. The model went beyond what we had ever done as humans."
「我们现在已经看到,在一月份曾有我不知道的 10 或 12 个公开的数学问题被解决,主要是由 GPT 5.2 解决的,现在最近有一些是由 Gemini 解决的。模型已经超越了人类知识的边界。我还不敢说它们正在解决人类无法解决的问题……模型超越了我们作为人类所做的任何事情。」

而这道"remix vs 真新颖"分歧最深的版本,来自 Lila 的 Andy Beam——他说仅靠强化学习"解题"不足以构成科学超智能,还需要主动"提出有趣问题"的开放式创造力(Ken Stanley 式的 open-endedness):

"You can't have scientific superintelligence if you're just a good test taker. So, like, if you think about what reinforcement learning is doing, even at scale, it's answering questions in kind of like a ruthlessly Vulcan-esque... you probably only in, like, limited ways would think of that model as being supremely creative."
「如果你只是一个优秀的考试者,那么你就无法拥有科学超级智能。所以,如果你考虑强化学习在做什么,即使在大规模下,它的回答问题方式就像是毫不留情的瓦肯式……你可能只会在有限的方式上将那个模型视为极具创造力。」

注意这场分歧的精确形状:几乎没有人真的认为今天的 AI 在"从零发明理论",连最乐观的 Wei、Weil 都把自己的战果限定为"极难的组合"或"人类本可做到、只是没做到"。真正的裂缝不在"AI 有没有拓展知识"(多数人认可它偶尔做到了),而在"这种拓展的量级"——是 Deedy 的小树枝,还是通往新枝干的第一步。这条线没有被任何两位发言人当面对质过。

二、纯推理 / test-time compute 能否替代新实验?

这是"实验护城河"共识背面的裂缝。Altman 把问题问得最纯粹——不建更大的加速器、不要更多数据,一个足够聪明的 AI 能否只凭现有数据把高能物理算出来?我们根本不知道智能在没有新实验时能走多远:

"I've always joked that one thing we should do when we have enough money ... is just build a gigantic particle accelerator and solve high-energy physics once and for all. ... But I wonder, what are the odds that a really, really smart AI could look at the data we currently have with no more data, no bigger particle accelerator, and just figure it out? ... we don't know how far intelligence can go with no more experiments. How much more could we figure out?"
「我一直开玩笑说,当我们有足够的钱……我们应该做的一件事就是建造一个巨大的粒子加速器,彻底解决高能物理问题。……但我想知道,一个非常非常聪明的人工智能,在没有更多数据的情况下……就能解决这个问题,这种可能性有多大?……我们不知道在没有更多实验的情况下,智能能走多远。我们还能弄清楚多少?」

他还给出一类不需要新实验的发现——已有数据里未被发掘的组合,比如已知的药物换个用途、稍加改动就接近重大突破:

"I suspect there's a lot of other examples that we'll find where Maybe we already have existing drugs that we know do something good, but they're reusable in some other big way or with a couple of small modifications, we are very close to something great. And it's been very heartening to hear from scientists using even the current generation models for this kind of work."
「我怀疑我们还会发现很多其他的例子,也许我们已经有现有的药物,我们知道它们可以做好事,但它们可以以某种其他重要的方式重复使用……听到科学家们使用即使是当前的模型进行这种工作,真是令人振奋。」

站在正对面的是 Krause,他认为反馈循环长是"根本性"瓶颈——数学能在几小时里跑完科学要几周几年的实验,这道坎绕不过去:

"the hardest part about AI for science is that our feedback loops are long, right? That's fundamental... you think about math, like AI and math, right? You can run a lot of experiments in hours that will take us weeks or years to run in science. How do you get around that problem? That's a really hard problem to solve."
「科学中人工智能最困难的部分是我们的反馈循环很长,对吧?这是根本性的问题……想想数学,像人工智能和数学,对吧?你可以在几个小时内进行很多实验,而我们在科学上需要几个星期或几年的时间才能完成。你怎么解决这个问题?这是一个非常难以解决的问题。」

Weil 把边界画在领域之间——数学物理可以纯 in silico,但从夸克到细胞到人体生物学的"第一性原理"模型还很远,实验重要:

"those things require labs. You can't do those just in silico. Although I think the importance of simulation is going to go up pretty meaningfully. We'll be able to apply huge amounts of compute to these problems, but then you're still going to need experimental validation. You're going to need to try things in the real world. I think it's going to be a while before we have a model that can … first principles go from like a quark all the way through a model of a cell all the way through human biology. Experiment matters."
「这些东西需要实验室。你无法仅仅在计算机上完成这些。虽然我认为模拟的重要性会有相当显著的提升……但你仍然需要实验验证。你需要在现实世界中尝试一些东西……从夸克出发,完整地模拟一个细胞的模型,一直到人类生物学。实验非常重要。」

这条分歧最微妙的地方在于:支持"纯推理能走很远"的最强证据(Erdős、IMO)全部来自数学——一个天然拥有闭环验证器、可完全 in silico 的领域。Krause 恰恰用数学来反衬材料的不同。所以 Altman 的乐观和 Krause 的护城河论,很可能根本不在谈同一类科学;纯推理能替代的,也许只是那些本就不需要新实验的领域。

三、需要海量数据吗?统一表征是不是必要条件?

第三条裂缝关于前提。数据规模上,Krause 直接顶回机器学习界的"你怎么拿到几百万数据点"——他说材料领域根本不需要,做新发现今天没见过这个要求(每种合金只产出约 50-150 个数据点):

"when I talk to the ML side, they're like, how are you going to get the millions of data points? And I'm like, I just don't think you need to. We have not seen that you need to, to do new discovery yet today."
「当我和机器学习方面的人谈话时,他们会问你怎么获得数百万的数据点?我就说,我认为你不需要。我们尚未看到今天需要这样做以实现新的发现。」

Lila 押在完全相反的方向——他们组装了 10 万亿个经实验验证的科学推理 token,把数据规模当核心引擎;但关键限定是,他们优先要"泛化性和灵活性"而非原始吞吐,要模型能设计连自己都没想过的新协议:

"We don't want to generate the same kind of data over and over again... the experimental platform that we're building prioritizes generalizability and flexibility over raw throughput. We want the model to be able to design a new experimental protocol, run the protocol, and receive the feedback, even if that's not an experiment we have thought about doing ourselves."
「我们不想一遍又一遍地生成相同类型的数据……实际上我们正在构建的实验平台优先考虑泛化性和灵活性,而不是原始的吞吐量。我们希望模型能够设计一个新的实验协议,运行协议并接收反馈,即使这不是我们自己想到的实验。」

表征上,Labenz 预言只有当模型对小分子、蛋白质、材料建立起统一表征、反馈开始来自现实之后,才可能出现类似超级智能的科学发现能力——统一表征是那道闸门:

"AI is not synonymous with language models. AI is being developed with pretty similar architectures for a wide range of different modalities, and there's a lot more data there. Feedback is starting to come from reality… When we start to give the next generation of the model these power tools, and they start to solve previously unsolved engineering problems, I think you start to have something that looks Kind of like superintelligence."
「人工智能并不等同于语言模型。人工智能的开发采用了非常相似的架构,适用于各种不同的模式,并且有更多的数据。反馈开始来自现实……当我们开始将这些强大的工具交给下一代模型,并且他们开始解决以前未解决的工程问题时,我认为你开始看到一些类似超级智能的东西。」

Lila 的 Rafa 直接反对这个前提——统一/语言化表征不是科学超智能的必要条件,很多数据模态与语言差别太大,未必值得都蒸馏进语言:

"I don't think that's a necessary condition for a scientific superintelligence... there was this quote from Demis Hassabis... it might not be worth distilling all the ways that live in sort of, you know, protein ester, alpha-4... there are data modalities that are so different from language."
「我认为这不是必要条件……可能没有必要提炼所有的方式,像是你知道的,蛋白质酯……或者还有一些数据模态。它们与语言是如此不同。」

而 Lila 自己的实证恰好倾向 Rafa——横跨生命科学、化学、材料的通用推理模型,常常跑赢领域专用模型,暗示存在一种可迁移的"科学推理"共性,即便各领域的"语言"表面上完全不同:

"So we have assembled this Reasoning data set of 10 trillion scientific tokens, reasoning traces that are experimentally verified across life sciences, chemistry, and material sciences. And we have seen that this general model often beats the domain-specific models."
「所以我们已经组建了这个包含 10 万亿科学标记的推理数据集,实验验证的推理轨迹涵盖生命科学,化学和材料科学。我们已经看到,这个通用模型通常优于特定领域模型。」

Rafa 还在方法论上补了一刀,切住"remix vs 真新颖"的边界——不能因为是 AI 就放松科学严谨标准,AI 产出的"惊喜"必须用与人类科学同等的验证标准来对待:

"We cannot relax our standards of scientific rigor because it's AI, right?... now we're past that, and we need to hold AI science to the same standard we hold regular human-led science."
「我会说我们不能放松科学严格标准,因为这是人工智能,对吧?……现在我们已经过去了,我们需要将人工智能科学保持在与常规人类主导科学相同的标准上。」

都没说透的

我的看法

以下是我的判断,把握程度中等偏上。

这九篇最大的发现是:表面的乐观共识比看上去窄得多。几乎所有人都同意 AI 已从纯工具跨进"偶发拓展前沿",但真正承重的命题——"AI 能不能真正发现新科学,而不只是高级 remix"——一追问就沿着定义裂开。而且裂缝的位置很反直觉:亲手攻下 Erdős 的 OpenAI 团队,反而是划界最保守的一方(Lijie Chen 的"组合已有想法 vs 从零生成"、Wei 的"发明新数学是数年到数十年"),而离一线最远的 Altman 反而最激进。这本身就说明,越接近产出、越清楚它的机理是"极难的搜索与组合",就越不敢称它为"新枝干"。

我以中高把握相信:AI 已跨过"能偶发拓展前沿"的门槛——最干净的证据不是 Erdős 反驳本身,而是它一周内外溢、被数学家用来反驳另一个关于实数的乘积猜想。孤立的成果可能是过拟合或巧合,能被独立迁移的思路才是真实的知识增量。同样,MIT 机制全新、对耐药菌有效的抗生素,很难用"人类花时间也能发现"完全解释掉。

但我也以中高把握判断:Deedy 的"小树枝 vs 新枝干"和 OpenAI 团队的"数年到数十年",比 Altman 的"5-10 年内 dwarf everything"更接近现实。今天所有拿得出的战果,无一例外是在人类已有的原语(class field theory、已知靶点、已知合金族)之上做极高难度的组合与搜索;"主动提出有趣问题"(Andy Beam 的 open-endedness)这一步没有任何实证,连 OpenAI 团队都承认 Erdős 那批问题是他们自己挑的子集。remix 能不能连续地长成"发明分支",是这个主题真正悬而未决的赌注。

而我认为最被低估、也最结构性的洞见,根本不在"模型多聪明"这一侧,而在 Krause 和 Lila 反复讲的"实验才是护城河"。约束 AI 科学速度的,在多数领域不是推理能力,而是物理世界的反馈通量——这解释了为什么数学、天体物理(能 in silico 闭环、或数据本就海量)先动,材料、生物(experiment-constrained、无统一表征、sim-to-real gap)殿后。领域进度不均不是暂时现象,而是被"数据从哪来、验证要多久"这条轴锁死的排序。顺着这条逻辑,Altman 那个"一千亿加速器 vs 纯推理啃现有数据"的问题,是这九个人里最重要、也最没被回答的一个赌注——它直接决定"纯推理替代新实验"这条捷径到底存不存在。

一个提醒把握度的注脚:本主题几乎全部由乐观阵营(三家 AI-for-science 创业公司 + OpenAI 系)构成,唯一的怀疑声音 Deedy Das 也只是"不可知论"而非反对。真正的空缺是一个"AI 做不了新科学"的硬派论证——没有它,这里的"共识"更像是一个自选样本,而非经过对抗检验的结论。

还想知道什么

取材

擦边未纳入(经挖掘判为 TANGENTIAL):