Evals 即护城河 · Evals as Moat
主题综述
更新日志
- 2026-08-24 — 新增 6 篇(Fireworks Lin Qiao、Mercor Foody-RL 环境、Trajectory Arjun、LangChain Harrison、Trajectory Ronak×2),全部纳入。图景三处实变:① 新开阵营 G「eval 的下一站是生产 trace——eval、训练、产品合流」:Trajectory 两位创始人 + Harrison 汇聚出「eval 取材自生产流量、纠正行为取代打分、trace 自动炼成 evals/judges/环境」这条线——08-23 随 Arena 信源退库而整块清空的「真实用户流量即 eval 底料」押注,由全新信源复活且更激进(不止喂 eval,直接喂训练)。② 阵营 F 加固的同时被劈成两条路线:Satya「私有 eval=IP、商品化模型层」获 Harrison 大会逐句转引(已成行业 talk track)、Foody 给出 12 个月时间表;但 Lin Qiao 给出反向路线——「把品味后训练进自有权重、无人可偷」,终点资产是模型权重而非 eval set,两条路线都叫 own your intelligence、对「模型层该不该商品化」的答案却相反(换模型自由 vs 养自家模型),正文新增张力段,Satya 换模型测试的适用域随之缩小。③ Foody 第三次出场改写阵营 B/E:eval 劳动生意升级为 RL 环境生意(worlds/apps/tasks 三件套,单任务 $50–$10K、前沿实验室月购 5 万任务,Mercor 四个月 $1B→$2B run rate),其「人类是测量前沿之外的唯一仪器 / 让模型自评=让学生自己改作业」从供给侧印证阵营 E 的 evaluator 瓶颈——并暴露 E 的瓶颈恰是 B 的收入来源这层利益结构。另:主流共识第一点补 Ronak 训练侧一票(benchmark 数月即饱和且与真实使用脱节);阵营 C 补一条覆盖面空白(职业任务 60-70% 涉人际协作、现有 eval 覆盖 <1%);阵营 D 补标量 reward 的算法侧证据(87/100 作文分类比),信号富度光谱三层成型。「都没说透的」新增两条:Camp F 路线分裂无人对质;本批六篇说话人全是闭环卖方、回声室风险。无未纳入篇目。
- 2026-08-23 — 退库清理:用户裁决"全退",移除 3 篇自动入库访谈的引用与相关论述(本篇 8 处)。图景有实质变化:阵营 C 失去第四种押注(Arena「真实用户流量即 eval 底料」整块删除),阵营 F 由四信源缩回三信源(Mercor/Satya/Nebius),「eval 即护城河以模型层竞争为条件」这一条件性前提随 Anastasios 信源一并撤出;Andreessen&Dixon 与 Sriram 两条本就在「未纳入」记录里,删除不改图景。
- 2026-08-22 — 回滚:当日自动入库的 17 篇访谈被用户否决、全部退库(Podwise Database 为手工精选库,自动 intake 已废除)。pipeline 补跑会话当日基于其中 6 篇对本篇的重写一并撤销,恢复至 2026-08-11 核验版(该版引用均已逐条核对)。
- 2026-08-10 — 新增 7 篇(Arena-Anastasios、Noam Brown、Lila Sciences、Nano Banana、Sriram、Baseten、a16z-CLARITY),其中 3 篇经逐字稿检查仅词面/邻接相关未纳入(a16z-CLARITY 系 crypto 立法纯误中;Sriram 谈 harness-moat,属 ai-moat-2026 议题;Baseten 是推理工程 QA)。纳入的 4 篇改了三处图景:① Noam Brown 给「主流共识第一点 + 阵营 C」补了 benchmark 的第三种垮法——能力已是测试时算力预算的函数,单数字 grid 是全行业心知肚明却无人敢先退出的「坏均衡」;其安全推论(危险能力上限=攻击者预算的函数)一并纳入,并在「都没说透的」新增:一切 eval 结论都是预算相对的,Satya 的换模型爬坡测试需加「同预算」条款。② Arena 的 Anastasios 从中立平台侧独立印证 Foody「eval 是部署最大瓶颈」(阵营 C 新增第四种押注:真实用户流量即 eval 底料),其「帮企业把自有 trace 炼成 eval、数据成为企业自己的护城河」+ $100M+ ARR 加固阵营 F;但他同时供出一个没人说过的前提——模型层若归一,eval 生意就死(eval 即护城河以模型层竞争为条件)。③ Lila Sciences 把阵营 E 的 evaluator 前提推进到物理世界:自然/实验=终极 verifier、思维链是不可靠叙述者、「实验验证过的推理 trace」是预训练语料里几乎不存在的稀缺资产——阵营 E 与 F 在此合流。Nano Banana 则属印证性补强、未改判断:阵营 A 的 eyeball 教义在前沿实验室内部同样成立(拿团队成员的脸做 eval set 逐一目检)。
- 2026-07-16 — 引用忠实度修复(全站审计):5 处——重引 3 条(Hamel"annotation and counting"、Kyle RULER/25→50、Shunyu Yao evaluator:均由 podwise 摘要句换回逐字原话)、更正归属 2 条("golden dataset sucks"系未具名共同主持人所说,换成 Aman 的逐字回应;"purest sense of PRD"系主持人 Lenny 所说,并补 Shreya 的逐字认可与限定)。
- 2026-06-11 — 取材升级为逐字稿全文。本次把 podwise 摘要层里的三手转述换成第一人称原话,重灾区是阵营 F(私有 eval = 护城河)——它原本三条引用全是 3rd-person 摘要,现已换成 Foody / Satya / Nebius 的逐字原话;并确认"Full-Stack Builder"那篇说话人就是 Satya Nadella。主流共识与阵营 B/C 的几条 paraphrase 也换成了逐字原话(Mia 的 SWE-Bench"saturated and contaminated"、Foody 的"eval = PRD = sales collateral"、median pay $95→$500)。新增逐字稿里被摘要层埋掉的料:企业"怕做 eval、因为会暴露自己被自动化"(Foody)、SWE-Bench 污染的具体证据(Mia)、"2026 是让模型端到端克隆 Slack 的一年"(Foody)。
- 2026-06-10 — 刷新:新增 3 篇(Brendan Foody/Mercor 新访谈、Satya Nadella、Nebius Roman Chernin),汇聚出"私有 eval set = 企业 system of record"这条新观点,新增阵营 F。
- 2026-05-20 — 首次综述。基于 8 篇访谈。
主流共识
第一点:通用公开 benchmark 在迅速饱和和污染——这一点全员同意,而且 OpenAI 自己的人讲得最狠。
"SWE-Bench Verified has been one of the North Star coding benchmarks that the field has looked at to measure coding progress. But recently, we've seen that progress has kind of stalled. And this we realize that this is because the eval is effectively saturated and also highly contaminated. So at this point, we think that it's not really measuring coding performance improvements well anymore."「SWE-Bench Verified 一直是全行业用来衡量编码进展的北极星基准之一。但最近我们看到进展停滞了。我们意识到,这是因为这个 eval 实际上已经饱和、而且被严重污染。所以到这个点上,我们认为它已经不能很好地衡量编码能力的提升了。」Mia Glaese / Olivia Watkins (OpenAI) · SWE-Bench-Dead
而且"污染"不是泛泛而谈——他们用一个审计 agent 抓到了具体证据:
"In SWE-Bench Verified, we found many instances of contamination across OpenAI models, across like Quad, Opus 4.5, Gemini, Flash. … We saw things like regurgitating the ground truth solutions, things like in some cases, giving like the task IDs …"「在 SWE-Bench Verified 上,我们在 OpenAI 的模型、Claude、Opus 4.5、Gemini、Flash 上都发现了大量污染实例。……我们看到模型直接背出 ground truth 答案,有些情况下甚至报出 task ID。」Mia Glaese / Olivia Watkins (OpenAI) · SWE-Bench-Dead
Galileo 的 Pratik Bhavsar 从另一个角度确认同一件事——通用榜单的排名不迁移到 agentic 任务:
"It's not necessary that what you see as top models in let's say LLM arena or other specific evaluation or general evaluations, they might not be also the same ranking for other tasks like agentic tasks."「在 LLM arena 或其他特定评估、通用评估上排名靠前的模型,未必在别的任务上、比如 agentic 任务上是同样的排名。」Pratik Bhavsar (Galileo) · Ranking Agentic LLMs
2026 年中,Noam Brown(OpenAI)补了第三种垮法——跟饱和、污染都无关:能力已经是测试时算力预算的函数,而单数字 benchmark grid 拿掉了这条 x 轴。5.5 发布时一度被 grid 误读成「没什么进步」,实情只是它思考得更省:
"my claim is the proper way to evaluate the models now is you either have some kind of Budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of test-time compute that's going into the model."「所以我认为,正确的评估模型的方法是,你要么有某种基准预算,无论是代币、成本、时间还是其他,要么你绘制出测试时间计算量与模型性能之间的关系图。」Noam Brown (OpenAI) · Why Traditional Benchmarks Fail
2026-08 再加训练侧的一票。Trajectory 的 Ronak Malde(前 Windsurf SWE-1 训练者,经 DeepMind)把饱和周期说得比谁都短,并在贵、慢之外补了第四个维度——与真实使用脱节:
"we are left with a bunch of benchmarks that are getting more time consuming, more and more expensive, and perhaps more concerningly, they're not tied to real world use cases where people are using AI."「我们面临着一堆基准,变得越来越耗时、越来越昂贵,或许更令人担忧的是,它们与现实世界的使用案例没有关联。」Ronak Malde (Trajectory) · Scaling up Continual Learning
第二点:LLM 实验室自己的 north star 已经是 eval-quality 而不是模型 capability——这把 evals 推到了基础设施级地位。Foody 把它压成一句话:
"If the model is the product, then the eval is the product requirement document. … In many ways, the barrier to applying agents to the entire economy to automate every workflow is how do we measure success? How do we eval it?"「如果模型是产品,那么 eval 就是产品需求文档(PRD)。……在很多意义上,把 agent 推向整个经济、自动化每个工作流的瓶颈,就是:我们怎么衡量成功?怎么 eval 它?」Brendan Foody (Mercor) · Why experts writing AI evals
第三点:evals 的真正信号在"看 trace"那一步——多个 practitioner 独立观察到。
"The right answer is keep looking at traces until you feel like you're not learning anything new."「正确的做法是持续看 trace,直到你感觉学不到新东西为止。」Hamel Husain & Shreya Shankar · Why AI evals are the hottest new skill
分歧在哪
阵营 A · "Evals 是 PM 的核心技能"——practitioner / 普及派
Aman Khan (Arize Head of Product) 给的是最 PM 化的版本:
"The PM's job is to have judgment on what that end product experience should be. And so being in the details on that when it comes to the human evals is really what determines whether or not your product is successful or fails."「PM 的工作就是对最终产品体验该是什么样做出判断。所以在人工评估上深入细节,是决定你产品成败的关键。」Aman Khan · Complete Beginner's Course on AI Evaluations
"golden dataset 决定后面一切"这个点,逐字稿里其实是共同主持人先说出口的(transcript 仅标注为 Speaker 3,通篇未具名):"if the golden data set sucks, then the rest of your evals will be terrible"。Aman 当场全盘接住,并把它钉成整个流程的第一步:
"Yeah, totally. … it's the most important step before you start trying to build anything complicated on top, like LLM as a judge or code-based evals. Like, just look at the data and debate, are these the right metrics for us to look at? Do we have the eval criteria in place? Do we know how to evaluate this or not? … we just start with five rows of data, right?"「是的,完全正确。……这是在你开始尝试构建任何复杂的东西之前最重要的一步,比如 LLM 作为裁判或基于代码的评估。比如,看看数据并争论,这些是我们需要关注的正确指标吗?我们是否有适当的评估标准?我们知道如何评估这个吗?……我们从五个数据行开始,对吧?」Aman Khan · Complete Beginner's Course on AI Evaluations
Hamel Husain 把同一立场推得更绝对:
"The most valuable process of evals is the annotation and counting. Even if that's all you do, you don't build any judge, you don't do any eval, you don't do whatever, you can get insane value by just doing that. That's the one part that everyone skips."「评估中最有价值的过程是注释和计数。即使你只做这些,你不构建任何 judge,不做任何评估,不做任何其他事情,仅仅通过做这些就能获得巨大的价值。这是每个人都会跳过的部分。」Hamel Husain · AI Evaluations Crash Course in 50 Minutes
"I've found that when you try to hide behind a score, you're not really making a decision."「我发现:当你试图躲在一个分数背后时,你其实没有在做决策。」Hamel Husain · AI Evaluations Crash Course
"eval-judge 即 PRD"这个说法,逐字稿里其实出自主持人 Lenny 的现场总结——他说最近多位嘉宾都在讲 "evals are the new PRDs",并指着屏幕上的 LLM-as-judge prompt 说:
"This is the purest sense of what a product requirements document should be. Eval, judge, that's telling you exactly what it should be and it's automatic and running constantly."「这就是 PRD 最纯粹的样子。Eval 和 judge 准确告诉你产品该是什么样,而且是自动的、持续运行的。」Lenny(主持人,综合多位嘉宾的 "evals are the new PRD" 说法) · Why AI evals are the hottest new skill
Shreya 当场认可,但补了一个 Camp A 式的限定——这份"PRD"必须从自己的数据里长出来,而不是事先拍脑袋写下:
"Yeah, absolutely. And it's kind of derived from our own data. So of course, it's a product manager's expectations. What I find a lot of people miss is they just put in what their expectations are before looking at their data. But as we look at our data, we uncover more expectations that we couldn't have dreamed up in the first place. And that ends up going into this prompt."「是的,当然。而且它在某种程度上来源于我们自己的数据。所以,这当然是产品经理的期望。我发现很多人忽略的是,他们在查看数据之前,就把自己的期望放进去了。但是当我们查看数据时,我们发现了更多最初无法想象的期望。而这些最终会进入到这个提示中。」Shreya Shankar · Why AI evals are the hottest new skill
注意——Hamel 本人在另一场访谈里说"躲在分数后面就是没做决策"。跟"eval-judge 即 PRD"放在一起看似矛盾,实际上区分的是自己理解后建立的 judge(PRD-级别)和外包给一个 generic metric 的躲避(hallucination score、coherence score 之类)——Shreya 的限定(judge 必须从自己的数据里长出来)恰好点在这条分界线上。
2026-08 刷新补一个前沿实验室内部的印证——「先用眼睛看」不只是应用层 PM 的穷办法。Google 的 Nano Banana(Gemini 2.5 Flash Image)团队开发新模型时的第一层 eval,就是拿团队成员自己的脸做 eval set、逐一目检:
"We've now created an eval set with that has, you know, my face in it and like other people on the team that we just kind of eyeball when we develop a new model."「我们现在创建了一个评估集,里面有我的脸,还有团队里其他人的脸,我们在开发新模型的时候会大致看看。」Nicole Brichtova (Google) · Nano Banana
在「判断权最终归谁」上,Oliver Wang 给了一个与 Hamel「别躲在分数后面」同构的回答——图像这种主观域,最好的 eval 是真实用户拿自己的 prompt 来裁决(transcript 把 LM Arena 转写成了 "El Marina",原样保留):
"But I think ultimately people are the arbiters of the images that they're trying to create for themselves. So that's why I think cases like El Marina where users are entering their own prompts, like this is the best way to evaluate a model."「但我认为最终人们是他们自己想要创作的图像的仲裁者。所以我才认为像 El Marina 这样的案例,用户输入他们自己的提示,是评估模型的最佳方式。」Oliver Wang (Google) · Nano Banana
他们也在跑 LLM-as-judge(用语言模型评自己的生成、形成 feedback loop),也会照着 X 上的社区反馈改 eval 以防倒退——但这两层都排在「自己人先目检」之后。
2026-08 补两笔。Fireworks 的 Lin Qiao 从基础设施侧把 Camp A 的教义翻译成工程语言——vibe eval 里的「感觉」其实就是创始人的判断,值钱的下一步是把它转成可重复的系统(transcript 将 systematic evaluation 转写为 "systemic evolution",原样保留):
"a lot of evals is vibe evaling. And the founders really look at the result and feel, hey, is this right or not? Actually, this is judgment. You put your judgment of any result and decide whether it's good or not. And that judgment you convert into systemic evolution. So this is no different from traditional software development, where you have your unit test, the integration test to ensure quality."「很多评估都是情感评价。创始人实际上会查看结果,感到,嘿,这是正确的吗?实际上,这就是判断。你对任何结果做出判断,并决定它是好还是不好。而这个判断转化为系统性的演进。这与传统软件开发没有区别,你有单元测试、集成测试来确保质量。」Lin Qiao (Fireworks) · Post-Training Is How You Keep Your Taste
她对判断权归属的说法与 Aman 完全一致:数据质量的最佳裁判是产品团队,后训练时代产品团队与 ML 团队正在合并成同一个组织。LangChain 的 Harrison Chase 则给「看 trace」补了一个机制性解释——为什么看 trace 比看输出有信息量:
"So when agents mess up, they mess up because an LLM call goes wrong. Why might it go wrong? It might go wrong for one of two reasons. One, the model is not good enough. Two, the context that the LLM received isn't good enough. And so I actually think it's the second one that more often than not causes issues."「当代理出错时,它们出错是因为 LLM 调用出现问题。为什么会出错?可能出错有两个原因。一是模型不够好。二是 LLM 收到的上下文不够好。所以我实际上认为,更多情况下是第二个原因导致问题。」Harrison Chase (LangChain) · When to Build Your Own Agent Harness
失败多半在 context 而非模型——所以只看最终输出永远诊断不出来,必须看 trace 里 context 是怎么一步步堆出来的。Hamel 的「看到不再学到新东西为止」,在这里有了因果论证。
阵营 B · "Evals 是经济级护城河 / 新职业"——Mercor 立场
Brendan Foody (Mercor) 把 evals 重新框架为一种新职业类别和供应链生意:
"We grew from one to 400 million in revenue run rate in 16 months, fastest ascent in history."「我们 16 个月内做到 $4 亿 ARR——史上最快增长。」Brendan Foody · Why experts writing AI evals
"If you really think about it, we were put on earth to create reinforcement learning training data for labs."「认真想想,我们生下来就是为了给实验室创造 RL 训练数据的。」Brendan Foody · Why experts writing AI evals
eval 同时是 PRD、也是销售材料——这条他讲得很完整:
"Evals are the PRD, but also subsequently the sales collateral, right? Because like evals are what you give to researchers to show them what they should be building … but they're also the way that you demonstrate the efficacy of capabilities."「Eval 是 PRD,但随后也是销售材料。因为 eval 是你给研究员看、告诉他们该构建什么的东西……但它同时也是你向外展示模型能力有多强的方式。」Brendan Foody · Why experts writing AI evals
Mercor 的经济结构本身是这条论点的实证——专家时薪远高于众包:
"Our median pay rate in the marketplace is $95 an hour, but it can flex up well up into, like, $500 an hour. … If you look at the economics of the crowdsourcing companies, oftentimes they would pay, like, $30 an hour to talent as sort of the average."「我们市场里的中位时薪是 $95/小时,但能上浮到 $500/小时。……而众包公司的经济结构,付给人才的平均值往往只有大约 $30/小时。」Brendan Foody · Why experts writing AI evals
Foody 在 Camp B 内部又给了一个独立判断——eval 业务的天花板取决于"人类还能做什么模型不能做的事":
"The market is bound by the amount of things where humans can do something that models can't."「这个市场的上限是:人类还能做、而模型做不到的事情的总量。」Brendan Foody · Why experts writing AI evals
逐字稿里挖出一条摘要层没有的料——Foody 观察到企业对"做 eval"本身的恐惧,因为 eval 会把"我正在被自动化"这件事坐实:
"There are certain enterprises we talk to that are almost fearful, not wanting to engage, not wanting to eval their businesses, because that'll provide the evidence that their value chain is being automated."「我们接触的一些企业,几乎是恐惧的——不愿参与、不愿对自己的业务做 eval,因为那会坐实'我的价值链正在被自动化'这个证据。」Brendan Foody · Why experts writing AI evals
2026-08,Foody 第三次出场把这门生意的形态又推进了一代:从「专家写 eval」到「专家造 RL 环境」——worlds(拟真语料环境:邮件、文档、表格)+ apps(Salesforce / ServiceNow / Microsoft 365 的高保真克隆)+ tasks(prompt + verifier)三件套。规模数字也换了量级,主持人开场即是:
"Mercor, I think you guys grew from a $1 to a $2 billion revenue run rate in the last four months or so."「Mercor,我想你们在过去四个月中,从十亿美元的营收增长到了二十亿美元的营收速度。」主持人(未具名) · RL Environments Explained
定价从劳动小时价换成了任务价:单任务 $50 到 $10,000,某些前沿实验室一个月买 5 万个任务,人才网络单季度产出 250 万专家小时。Foody 给这门生意的战略定位也升级了:
"… the three core pillars of their AI strategy are their compute, their algorithms or researchers and the data sets they build. And data is often the most differentiating factor."「……其人工智能战略的三个核心支柱是计算能力、算法或研究人员以及他们构建的数据集。而数据通常是最具差异化的因素。」Brendan Foody (Mercor) · RL Environments Explained
注意他对 Camp B 天花板的说法也在进化:旧访谈说「市场上限=人类能做而模型不能做的事的总量」,这次他把人类的角色钉在一个更耐久的位置——不是劳动力,而是测量前沿之外的唯一仪器(这条与阵营 E 汇合,引语见彼处)。
阵营 C · "通用 benchmark 在垮,接什么各有押注"——OpenAI / Galileo 立场
OpenAI 的 Mia Glaese / Olivia Watkins 押的是"更难的公开 benchmark + 正确性之外的维度"。光看对不对已经不够了:
"Like Olivia talked about sort of like, does it have like design taste, right? Like does it solve the problem the way that, you know, my team likes to solve problems. Is the code nice, right? Like, is it, is it well written? Is it sort of like clean code, right? Like people care about this. Is it maintainable in the future? People care about a lot of these maybe less tangible, less tangible and like harder to measure, frankly, things that are still like super meaningful for people that are working with coding agents."「就像 Olivia 说的——它有没有设计品味?它解决问题的方式,是不是我团队喜欢的方式?代码漂不漂亮?写得规不规范?是不是整洁的代码?将来好不好维护?这些更难量化的东西,大家其实很在意,而且对用 coding agent 的人来说意义重大。」Mia Glaese / Olivia Watkins (OpenAI) · SWE-Bench-Dead
Pratik Bhavsar (Galileo) 押的是 domain-specific leaderboard 是下一站:
"People in the industry come from, let's say, I'm from healthcare. I want to see if this models work for healthcare or not. People are coming from investment, finance, insurance industry, and they want to know that is, are the models ready for my domain or not?"「业内的人会说,我来自医疗——我想看模型在医疗上行不行。来自投资、金融、保险的人也想知道模型在他们的领域上 ready 没 ready。」Pratik Bhavsar (Galileo) · Ranking Agentic LLMs
而且他给了一个具体的"饱和"数字——连他自己的新 benchmark 都已经被刷到 0.95:
"We saw with V1 that the scores are saturating. We see that the best score has already reached 0.95. … So we want the benchmarks to be harder."「我们在 V1 上就看到分数在饱和——最好的分已经到 0.95。……所以我们想把 benchmark 做得更难。」Pratik Bhavsar (Galileo) · Ranking Agentic LLMs
Noam Brown (OpenAI) 押的既不是更难的题、也不是垂直榜单,而是换坐标轴——一切评测画成算力预算的函数。他把「明知 grid 有误导还继续发」点破成一个博弈困境:
"People expect us to publish the grid. And then, okay, well, why do people expect the grid to be published? Because everybody publishes the grid. And so you kind of end up in this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out."「人们希望我们发布网格。那么,好吧,为什么人们期望发布网格?因为每个人都发布网格。所以你最终会陷入这种糟糕的平衡中,每个人都知道这是个糟糕的平衡,但没有人想要打破它。」Noam Brown (OpenAI) · Why Traditional Benchmarks Fail
这条对安全评估的推论更狠——「危险能力上限」根本不是模型常数,而是攻击者预算的函数(他引 AISI 的网络安全评估:跑到 1 亿 token 模型还在继续涨分):
"The problem is we're in a world now where The capability of the model is a function of how much money you put into it, basically. … At what budget should you evaluate these models?"「问题是我们现在处在一个模型能力基本上是你投入多少钱的函数的世界。……你应该在什么预算下评估这些模型?」Noam Brown (OpenAI) · Why Traditional Benchmarks Fail
2026-08 补一条「垮法之外的空白」。Foody 指出现有 eval 体系结构性漏掉了职业工作的主体部分——人际协作:
"my favorite questions to ask people when they're thinking about their data distribution is what percentage of tasks that they do in their job Require interacting with other people. And most people would say like 60 or 70%. … But then if you map that on to what percentage of evals measure how well the models can interact with other people, it's like 1%."「我最喜欢问人们关于他们数据分布的问题是,他们工作中需要与他人互动的任务占百分之几。大多数人会说大约 60% 或 70%。……但如果您将其映射到评估中,评估模型与其他人互动的能力占的百分比,大约是 1%。」Brendan Foody (Mercor) · RL Environments Explained
这不是饱和、污染或预算轴的问题,而是覆盖面问题:三种押注(更难的公开题 / 垂直榜 / 私有 eval)目前全都没碰它。Foody 顺带给出他押的下一站——超长时程任务(当前 agent 大多没训过 10 小时以上的任务,要往人类 100 甚至 1000 小时量级造题)+ 虚拟同事(把社交协作放进 eval)。
阵营 D · "LLM-as-judge 要分相对排序和绝对打分"——OpenPipe 的拐点
Kyle Corbitt (OpenPipe) 给了 evals 工具链里一个被许多人忽略的实证拐点:
"And so simplifying a lot, it's basically just LLM as judge on a whole group. So you say, okay, this is the task I'm trying to achieve. Here's four different runs of an agent trying to achieve it. Which of these did best? And it stack ranks them. And it turns out that works phenomenally well with gRPO, like way better than I expected …"「所以简化一下,它基本上只是把 LLM 作为对整个群体的评判。所以你说,好的,这是我试图完成的任务。这是代理尝试完成它的四种不同的运行方式。其中哪个做得最好?它对它们进行堆叠排名。事实证明,这与 gRPO 配合得非常好,比我预期的好得多……」Kyle Corbitt(谈 RULER) · Why Fine-Tuning Lost and RL Won
这个"相对排名"拐点直接改写了他自己对 RL 方向的下注概率:
"So I mentioned my initial opinion of how likely this direction was to work was maybe 25%. We're up to 55% or so. And RULER is actually a big update that got me from the 25 to the 50."「所以我提到了我最初的看法,即这个方向成功的可能性大约是 25%。我们现在达到了 55% 左右。而 RULER 实际上是一个很大的更新,它使我从 25% 提高到了 50%。」Kyle Corbitt · Why Fine-Tuning Lost and RL Won
"If you tell a human, choose which of these is better, it's easier for them to do than say, is this one good or bad in absolute terms?"「让一个人选'两者哪个更好',比让他说'这个绝对意义上好不好'要容易得多。」Kyle Corbitt · Why Fine-Tuning Lost and RL Won
跟 Hamel"看到 agreement score 就警惕"放一起,这揭示了一个没人正面说出来的区分——LLM-as-judge 在相对排名任务上(RULER)有效,在绝对打分任务上(generic hallucination/coherence score)容易自欺。
2026-08,Trajectory 的 Ronak 从训练算法侧给这条线补了地基:标量分数不只是「躲在分数后面」的决策问题,它在信息量上就是坏的——
"it's almost like you're drinking through a straw in order to get the reward. The way to think about it is imagine you were trying to write an essay and your teacher just gave you a score of 87 out of 100. You'd have to run through so many different examples to get to the idea of what a good essay is …"「几乎就像你在用吸管喝东西来获取奖励。这样想:想象一下你正在写一篇论文,而你的老师仅仅给了你一个 87 分(满分 100)。你必须经历这么多不同的示例,才能让你明白什么是好论文……」Ronak Malde (Trajectory) · Scaling up Continual Learning
他的解法(OPSD / SDPO)是把「纠正文本」当逐 token 的密集反馈,而不是把一切压成序列级的一个数字。这把 Kyle 的「相对排序 > 绝对打分」推进成一个更完整的层级:绝对单分 < 相对排序 < 逐 token 的文本级纠正——eval 信号的富度光谱,三段各有主人。而且 Hamel 的 annotation(产品层)、Arjun 的 corrective behavior(数据层,见阵营 G)、Ronak 的 per-token distillation(算法层),其实是同一个判断在三个层面的独立出现。顺带一条对称性:RL 的 reward hacking 在这条新路线里也有等价物——hint leakage(模型直接把提示里的答案抄进推理),eval 信号越富、被 hack 的面也越大。
阵营 E · "好的 evaluator 是 reflection / self-correction 的前提"——研究侧视角
Shunyu Yao 和 Harrison Chase 在讨论 reflection / self-correction 时给的限定值得拎出来:
"I think a key bottleneck is the evaluator … in order to reflect upon your thoughts, you have to have a very good evaluator to judge whether your thought is good or not. But that might be as hard as solving the problem itself or even harder. The principle of self-reflection is probably more applicable if you have a good evaluator, for example, in the case of coding. If you have those arrows [errors], then you can just reflect on that and how to solve the bug and stuff."「我认为一个关键瓶颈是评估器……为了反思你的想法,你必须有一个非常好的评估器来判断你的想法是好还是不好。但这可能和解决问题本身一样难,甚至更难。如果你有一个好的评估器,例如在编码的情况下,那么自我反思的原则可能更适用。如果你有这些箭头(transcript 原文如此,应为 errors),那么你就可以反思它以及如何解决 bug 等等。」Shunyu Yao · Language Agents: From Reasoning to Acting
这把 evals 从"产品工具"升级到"agentic 智能的前提条件"——没有好的 evaluator,多数 agentic 自改进都失效。
2026-07 的 Lila Sciences(Andy Beam & Rafa Gómez-Bombarelli)把这条线推到物理世界的尽头——整个公司的论题就是「把自然当 verifier」,在互联网语料耗尽之后用实验反馈当下一个数据源:
"At Lila, what we believe is that actually science, running the scientific method and using nature and experiments as verifier is like the ultimate version of that. … They are scaled verifiers for science so that we can do post-training at scale and push out the frontier of what reasoning models are capable of."「在 Lila,我们相信,实际上,科学、运行科学方法以及使用自然和实验作为验证者就像是这一切的终极版本。……它们(AI 科学工厂)是科学的规模化验证器,以便我们能够进行大规模的后训练,并推动推理模型的能力边界。」Andy Beam (Lila Sciences) · The Lab of the Future
比 Shunyu 更进一步的是「拿什么当 ground truth」的取舍——连模型自己的思维链都不可信,可信的只有 verifier:
"the chain of thought is often an unreliable narrator for what the model, the computation of the model is actually doing. … How much should we rely on the chain of thought versus just trusting the experiment, trusting the verifier, trusting the simulator as the ultimate ground truth?"「思维链通常是关于模型实际计算的一个不可靠叙述者。……我们应该在多大程度上依赖思维链,而不是仅仅相信实验、相信验证者、相信模拟器作为最终的真实依据?」Andy Beam (Lila Sciences) · The Lab of the Future
而这条研究侧的线在 Lila 身上直接接通了阵营 F——「实验验证过的推理 trace」正是预训练语料里几乎不存在、竞品爬不到的资产(他们已攒到 10 万亿 token 量级):
"If you think about an experimentally verified reasoning trace, how many of those do you think exist on the internet or in the pre-training corpus? We have just seen incredible lift from showing the model that even if we're at like a parameter disadvantage relative to the frontier models, just showing it in experimentally verified reasoning trace, you see just immediate lift when we do that."「如果考虑一下经过实验验证的推理痕迹,你认为在互联网上或预训练语料库中有多少?我们已经看到通过向模型展示,即使我们在参数上相较于前沿模型处于劣势,仅是展示经过实验验证的推理痕迹,当我们这样做时,你会看到立即的提升。」Andy Beam (Lila Sciences) · The Lab of the Future
2026-08,Foody 从数据供给侧给了 evaluator 瓶颈最直白的一版——为什么 RL 环境的 verifier 必须由人来写:
"But in most domains, like building a slide deck, The model has an incredibly hard time identifying reliably where it made its own mistake. It's as if you would be asking a human to grade their own homework."「但在大多数领域,如制作幻灯片,模型在可靠识别自己犯错的地方时会遇到极大的困难。就像你在要求一个人给自己的作业打分一样。」Brendan Foody (Mercor) · RL Environments Explained
"But the reason that humans are still an essential component of the process that's incredibly differentiated is that you need humans almost definitionally to measure what is beyond the frontier of the model capabilities."「但人类之所以仍然是这个过程的一个至关重要的组成部分,且具有独特性,是因为几乎可以定义为您需要人类来衡量超出模型能力边界的东西。」Brendan Foody (Mercor) · RL Environments Explained
这与 Shunyu「evaluator 可能比解题本身更难」完全同构,但推导方向相反:Shunyu 从研究侧推出「没有好 evaluator,自反思失效」;Foody 从生意侧推出「正因为模型自评失效,人类专家的测量权是 Mercor 商业模式的地基」。也就是说,阵营 E 的瓶颈恰好是阵营 B 的收入来源——这层利益结构,读他的每一句「人类不可替代」时都要带着。(他自己也划了例外:数学、网络攻防这类有干净模拟器/攻防对抗结构的域,verifier 可以不靠人。)
阵营 F · "私有 eval set = 企业的 system of record / 真正的护城河"——Mercor / Satya / Nebius 的汇聚
这是 2026 年 6 月这批访谈里最大的新观点,来自三个独立信源。逐字稿读下来,这条线比摘要层呈现得更硬、也更可证伪。
Foody 在新访谈里把 eval set 直接定义成企业的新 system of record,并点出它的战略用途——商品化模型层:
"Over time, this is going to develop to look very similar across every Fortune 500, where they'll need to have this system of record for evaluating and specifying agent behavior across every workflow in their business. And they're going to use that to commoditize the model layer because they want to enable perfect competition for the models having zero switching costs."「随时间推移,这会演变成每个 Fortune 500 都长一个样:他们需要一个 system of record,去评估和规定每个业务工作流里的 agent 行为。而且他们会用它来把模型层商品化——因为他们想让模型之间完全竞争、切换成本为零。」Brendan Foody (Mercor) · 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility
护城河不在软件层,而在 forward-deployed 把组织隐性知识编码进 agent:
"If you have a great forward deployed motion where you're going deep with a customer, you're training the agents based on all of this tacit knowledge within the company so that it understands how to perform effectively, that feels incredibly differentiated and hard to recreate."「如果你有一套很强的 forward-deployed 动作——深入一个客户、用公司里所有那些隐性知识去训练 agent,让它知道怎么高效干活——那才是极难复制、极具差异化的东西。」Brendan Foody (Mercor) · 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility
而软件层为什么没护城河——他给了一个吓人的时间表:
"We're building out an eval set that measures how effectively agents can build end-to-end SaaS applications, where 2025 was the year of, how do you get a model to make a PR and a code base? And 2026 is the year of, how do you get the model to clone Slack end-to-end? Those capabilities are going to exist in the models in the next 12 months."「我们正在做一个 eval set,衡量 agent 端到端构建 SaaS 应用的能力。2025 年的主题是'怎么让模型在代码库里提一个 PR',2026 年的主题是'怎么让模型端到端克隆一个 Slack'。这些能力,未来 12 个月内就会出现在模型里。」Brendan Foody (Mercor) · 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility
Satya Nadella 给了几乎一样的判断,但加了一个可操作的"控制权测试"——能不能在不泄露 trace 的前提下,把私有 eval 从模型 A 搬到模型 B 继续爬坡:
"Every company having private evals may be the biggest IP. … You have an eval that's private, you're using a model A, can you switch it to model B and climb up? If you can, then you're in control. If you can't, you're not in control."「每家公司拥有自己的私有 eval,可能就是最大的 IP。……你有一个私有 eval,现在用模型 A——你能不能换成模型 B 并继续往上爬?如果能,你就握有控制权;如果不能,你就没有。」Satya Nadella (Microsoft) · The Rise of the Full-Stack Builder
他还把"公开 eval 都能被刷爆、所以只有私有 eval 才算数"讲明白了:
"You'll have private evals because we know all the evals out there are good, interesting, but they're not really that critical at this point because they all can be maxed. And so the point is each company will have its own private eval."「你会有自己的私有 eval,因为我们都知道外面那些 eval 虽然好、虽然有意思,但此刻已经不那么关键了——它们都能被刷满。所以重点是:每家公司都会有自己的私有 eval。」Satya Nadella (Microsoft) · The Rise of the Full-Stack Builder
Nebius 的 Roman Chernin 从基础设施侧给了第三个佐证——eval 是企业能不能进入指数增长期的冷启动门槛:
"First of all, they were focusing on evaluations. … You need to have like metrics. … You need to have this CI, CD process established for AI development. … They have this, you can call it foundational investments, a cold start problem, how to start shipping. When they solve it, they start to grow exponentially."「首先,他们聚焦的是评估。……你得有指标,你得为 AI 开发建立起 CI/CD 流程。……这是一种基础投资,可以叫它'冷启动问题'——怎么开始交付。一旦解决,他们就开始指数级增长。」Roman Chernin (Nebius) · 20VC: Nebius Co-Founder on AI Infrastructure Bubbles
2026-08 刷新给这条线加了三块料,同时把它劈成了两条路线。
第一块:Satya 的 framing 已经成为行业 talk track。LangChain 的 Harrison Chase 在讲 harness 的大会演讲里,逐句转引 Satya 的文章当自己的论纲:
"One, create your private evals because eval defines what good looks like inside the organization. Two, retain ownership of your organization's memory, traces, feedback, … decisions and institutional context. And then three, you create your own continuous learning loop, hill climbing machine that will allow your AI investments to compound the value of your firm."「第一,创建你的私人评估,因为评估定义了组织内部的良好标准。第二,保持对你组织的记忆、痕迹、反馈、……决策和机构背景的所有权。第三,你要创建自己的持续学习循环、爬山机器,确保你的 AI 投资能够复合你公司的价值。」Harrison Chase(现场逐句转引 Satya Nadella 的文章) · When to Build Your Own Agent Harness
并给出自己的行业观察:
"I think every company, when they're building a mission critical agent, they will build benchmarks for that agent."「我认为每家公司在构建任务关键型代理时,都会为该代理建立基准。」Harrison Chase (LangChain) · When to Build Your Own Agent Harness
第二块:Foody 第三次出场,把「own your own intelligence」的时间表压到 12 个月内(transcript 把 moats 转写成 modes,原样保留):
"And I believe that over the next 12 months, there's going to be dozens of examples just like that, where companies own their own intelligence. And that is the key source of the modes that they're building."「我相信在接下来的 12 个月中,会有十几个这样的例子,企业拥有自己的智能。这就是他们构建模式的关键来源。」Brendan Foody (Mercor) · RL Environments Explained
第三块是真正的复杂化——Fireworks 的 Lin Qiao 给出了一条反向路线。她完全同意前提(应用层易克隆、必须 own 判断与品味),但她的终点资产不是 eval set,而是后训练进自有权重的模型:
"post-training becomes a very appealing solution because post-training allows you to basically encode, codify your unique taste Into a model that no one can steal from. Because it's very easy to clone and copy application as is, as you all know, right?"「后训练成为一个非常吸引人的解决方案,因为后训练让你基本上能够将你独特的品味编码成一个无人可以窃取的模型。因为像大家所知道的那样,克隆和复制应用程序是非常简单的,对吧?」Lin Qiao (Fireworks) · Post-Training Is How You Keep Your Taste
在她的版本里,eval / reward 不是独立的护城河资产,而是把品味写进权重的传动装置——reward 就是代码化的 rubric,维度配比即 secret sauce:
"rewards actually is code. You write rewards in code. And you should think about rubrics of rewards and think about you want to grade the result in multiple dimensions. … different company have a different way to blend those. And that's that's your unique part and secret sauce."「奖励可以认为是代码。你在代码中写奖励。你应该考虑奖励的标准,并考虑你希望在多个维度上对结果进行评分。……不同的公司有不同的方式来融合这些。这就是你的独特部分和秘密武器。」Lin Qiao (Fireworks) · Post-Training Is How You Keep Your Taste
这条传动装置的工艺难度她也给了第一手样本——reward 写歪一个维度,模型立刻钻空子:
"So a fun story about reward hacking is we have been asking a model to generate, this is a coding example, … To generate code that minimized the compilation error. So guess what the model did? The model generate zero line of code. Okay, there's no compilation error, but that's absolutely not what you want."「关于奖励黑客的一个有趣故事是,我们一直在要求一个模型生成,这是一个编码示例,……生成最小化编译错误的代码。所以你猜模型做了什么?模型生成了零行代码。好吧,没有编译错误,但这绝对不是你想要的。」Lin Qiao (Fireworks) · Post-Training Is How You Keep Your Taste
两条路线的张力没人点破:Satya/Foody 路线的要义是「eval 私有、模型商品化、零切换成本」——eval set 独立于任何模型存在,换模型自由正是控制权的证明;Lin Qiao(以及阵营 G 的 Trajectory)路线的要义是「品味蒸馏进自有权重」——资产与某一份具体权重焊死,「换模型能否继续爬坡」这个测试在这条路线里几乎失义(你不换模型,你养模型)。两条路线都自称 own your intelligence,对「模型层该不该商品化」给出的答案却正好相反。Lin Qiao 还给了自己路线的时机条款:PMF 之前不要碰后训练——只有 PMF 之后从产品表面收上来的数据才有意义,那才是「own 智能」的燃料;这与 Nebius 的「eval 是冷启动门槛」拼在一起,正好是同一条时间线的前后两段。
这条线和 Camp A/B 的关系微妙:Camp A 说 eval 是 PM 技能、Camp B 说 eval 是高薪专家生意——Camp F 说积累出来的那套 eval set 本身是护城河。三者其实是 eval 价值链的三个环节:写它的人(A)、外包它的生意(B)、攒成的资产(F)。但没人正面把这三环串起来。它还直接挑战姊妹主题 ai-moat-2026 / ai-native-products——如果"软件层没有 defensibility、私有 eval 才是护城河"(Foody 新访谈标题主张),护城河的定义又被重画了一次。2026-08 之后这条价值链还得再加两环:把 eval 资产转成自有权重的训练层(Lin Qiao / Trajectory),和把整个循环自动化的工具层(LangChain)——五个环节各有人占位,没人拥有全链。
阵营 G · "Eval 的下一站是生产 trace——eval、训练、产品合流"——持续学习派(Trajectory / LangChain)
2026-08 这批访谈里最大的新线,三个信源(Trajectory 两位创始人 + LangChain 的 Harrison Chase,Fireworks 的 Lin Qiao 从后训练侧呼应)汇聚出同一主张:eval 不再是独立环节——它的底料是生产环境里的真实 trace,它的终点是直接喂进训练循环。值得注明:08-23 Arena 信源退库时,「真实用户流量即 eval 底料」这条押注曾整块清空——现在它被全新的信源复活,而且更激进:不止拿流量当 eval,直接拿流量当训练信号。
Ronak Malde(前 Windsurf SWE-1 训练者、经 DeepMind,放弃收购分成创办 Trajectory)的前提判断是:静态模型在系统性浪费生产信号——
"the model that you used yesterday is going to be the same model and making the same mistakes tomorrow. And all of those corrections you gave it, the edits, like in any product is all just being put to waste."「你昨天使用的模型明天仍然会是同一个模型,并犯同样的错误。你给它的所有修正、编辑,像任何产品一样,都是浪费掉的。」Ronak Malde (Trajectory) · Every product of the future will be a living system
浪费的规模他给了个名字——"trillion token problem":
"we are actually spending hundreds of trillions of tokens every single day on inference. And we're generating great amounts of data on how models in the real world are failing, how they're doing well, and that should be signal that we should be capturing and training on."「我们实际上每天在推理上花费数百万亿代币。而且我们正在生成大量的数据,关于模型在现实世界中如何失败、表现良好,这应该是我们应该捕捉和训练的信号。」Ronak Malde (Trajectory) · Scaling up Continual Learning
联合创始人 Arjun Karanam 把「eval 从哪来」直接答成:从流量里来——
"this means that evals are, you know, drawn from traffic, how people are actually using your product, both how they're using it now and the things that they're requesting on the frontier that might not be possible, all super helpful."「这意味着评估是基于流量的,由人们实际使用你的产品产生的,无论是他们现在的使用情况还是他们在边界上请求的可能不可能的事情,都非常有帮助。」Arjun Karanam (Trajectory) · Continual Learning
信号形态上,两人同时点破 thumbs up/down 是坏 eval 信号,真正的 ground truth 是纠正行为:
"And that sounds amazing in theory, but it's incredibly noisy. And if you've used any coding agent, you know, you kind of just like accept everything that the agent does. And it's only like five commits later that you're like, oh, crap, like, This broke everything. Let me go and undo that. And so it's the corrective behavior, the edits, the undo's and the retries that both need to be elicited from the user, but also captured."「这在理论上听起来很棒,但实际上噪声非常大。如果你使用过任何编码智能体,你知道你基本上就是接受智能体所做的一切。只有在五次提交后你才会意识到,哦,糟糕,像是,这破坏了一切。让我去撤回它。所以需要从用户那里引导出校正行为、编辑、撤消和重试,但也要将其捕捉到。」Arjun Karanam (Trajectory) · Continual Learning
Trajectory 的产品主张就是把这些 trace 炼成统一格式,再自动生成 eval 的全套组件:
"We essentially take all of the data, all of the expert traces, the agents, And distill it into one format, which is what we call the trajectory. And it's basically all of the data that is needed to then create the evals, the judges, the environment, everything, all the components that you would need for training."「我们基本上把所有数据、所有专家痕迹和代理集中提炼成一种格式,我们称之为轨迹。基本上,这些数据是创建评估、判断、环境所需的所有内容,以及所有培训所需的所有组件。」Ronak Malde (Trajectory) · Every product of the future will be a living system
为什么这在专业域是生死线(与 Camp F 的 forward-deployed 论同一逻辑):
"For a field like legal, like getting 80% of the way there is the same thing as zero."「对于法律这样的领域,达到 80% 的结果与 0% 的结果是一样的。」Ronak Malde (Trajectory) · Every product of the future will be a living system
Harrison Chase 从工具链侧给出同构的 flywheel——跑 agent → 收 trace → curate → 实验 → 改 harness/模型/context,且已在用 agent 自动化这个循环本身(LangSmith Engine:一个盯着生产 trace 找失败模式、直接给 prompt/context/harness 提交修复建议的 agent,已经 Engine-on-Engine 自举)。他明确说这个 flywheel 就是 Satya 那三条的落地实现——阵营 G 是阵营 F 的执行层。
对本主题的冲击在于:如果 eval 的底料是生产 trace、eval 与训练合流,那么「私有 eval set」这个静态资产的概念本身开始溶解——护城河从「攒了多大的 eval set」滑向「trace→eval→权重这条循环转得多快」。Arjun 的判词是 eval、产品、训练三者在理想世界里应该是同一个东西("In an ideal world, these are all the same")。不过要带着读的是:本阵营(连同呼应者)全部是卖这个循环某一段的 vendor——Trajectory 卖训练层、LangChain 卖观测与工具层、Fireworks 卖算力与后训练平台——回声室风险记在「都没说透的」。
都没说透的
- "Evals 是普及技能(Camp A)"和"Evals 是专家高薪生意(Camp B)"如何并存? Hamel 说"PM 自己做",Foody 说"我们雇 $500/小时的专家"——没人正面解释这两条经济逻辑能不能同时成立。最可能的答案是不同 evals(产品级 vs 模型 frontier 级)适配不同人,但语料里没人讲透。
- LLM-as-judge 用作"相对排序"和"绝对打分"效果差很多,但几乎没人正面区分。 Kyle 用 RULER(相对)就工作了;Hamel 警惕的是 agreement/coherence(绝对)。没人画出"什么任务用哪种 judge"的决策表。
- 公开 benchmark 死后接什么?三种押注各异,没人 head-to-head。 OpenAI 押 SWE-Bench Pro(更难的公开 benchmark),Galileo 押 domain-specific leaderboard(垂直公开),Hamel/Aman/Satya 押完全私有的 internal eval。三条路径的相对优势没人正面比较。
- "私有 eval = 护城河"(Camp F)和"data moat 只在 mega-scale 成立"(a16z, 见 ai-moat-2026)正面打架。 Foody/Satya 说一套攒出来的私有 eval 就能护城;a16z 说数据网络效应要到数十亿用户才显现。私有 eval 到底要攒到多大、多久才真正难以复制?Satya 给了"能不能换模型继续爬坡"这个定性测试,但没人给定量门槛——这是 Camp F 论点最关键、也最空的一块。
- Foody 同一个人,两次访谈 framing 变了。 旧访谈卖的是"专家写 eval 的劳动生意"(Mercor 的 GMV、$500/小时),新访谈卖的是"私有 eval = 护城河 + 应用层没有 defensibility"。两个 framing 服务不同叙事(一个抬高劳动供给价值,一个抬高 Mercor 作为护城河中介的价值),没人追问他这两套说法怎么自洽——尤其是:如果护城河在客户自己攒的私有 eval 里,那 Mercor 这个"帮你攒"的中介,长期价值捕获在哪?(2026-08 更新:第三次出场给出第三个 framing——RL 环境生意,且对"中介价值捕获在哪"有了答案雏形:给关键客户配全独占的 siloed 团队 + 环境基建的规模经济,"看 Frontier Labs 都不自建就知道"。framing 三连跳每次都跟着 Mercor 当期产品走,追问依然成立,但至少这次他正面回答了中介的存在理由。)
- 一切 eval 结论都是预算相对的,但没人给自己的论点加预算条款。 Noam Brown 证明能力=测试时算力预算的函数(AISI 的 cyber eval 跑到 1 亿 token 还在涨),那么私有 eval 的爬坡曲线、domain leaderboard 的名次全都随预算漂移。Satya 的「换模型能否继续爬坡」严格说要补一句「在同一推理预算下」才成立——Camp F 没人处理这个变量。
- Camp F 的路线分裂:eval-set-as-IP vs taste-in-weights,没人对质。 Satya 的控制权测试以「换模型自由」为纲——eval set 是独立于任何模型的资产,模型层被商品化;Lin Qiao / Trajectory 的路线把品味蒸馏进自有权重——天然放弃换模型自由(或者说把模型本身变成自家资产,测试失义)。哪条护城河更耐久?eval set 可以带着走,但攒的方法人人可学;权重偷不走,但绑死基座模型与算力供给(Fireworks、NVIDIA 恰好是卖这两样的)。语料里两边各说各话,没有一场正面交锋。
- 「私有 eval / own your intelligence」叙事出现回声室结构。 Satya 的文章被 Harrison 在大会逐句转引、Foody 给出 12 个月时间表、Lin Qiao/Trajectory/LangChain 各自把它写进产品叙事——本轮新增六篇的说话人全部是卖这个循环某一段的 vendor(数据、训练、算力、工具),连一个付钱攒 eval / 训模型的买方第一人称都没有。共识在变厚,独立证据没有变多——话术趋同本身不证伪论点,但它让「vendor 合唱」与「事实收敛」变得无法区分。
判断(不是事实):evals 在 2026 年同时是两个东西——对应用层 PM,它是手工 + 判断密集的核心 craft(Camp A 对);对前沿实验室,它是新型供应链业务(Camp B / Mercor 对)。这两件事的混淆制造了大量错觉——很多公司既不严肃做手工 trace 分析、又付不起 Mercor 那种专家费,结果两边都不在。真正的护城河不在工具或方法(会快速商品化),而在积累专有 trace、并把它转成内部 judge 的能力——Satya 的"能不能换模型继续爬坡"是我目前见过最干净的判定标准。读完逐字稿我对 Camp F 的把握升了一点:因为 Satya 给了可操作测试、Foody 给了"克隆 Slack"的时间表,这条线不再只是 vendor 口号。
2026-08-10 补一笔:Noam Brown 逼我给 Satya 那个判定标准加一个条款:换模型爬坡必须在同一算力预算下比较才算数——否则「爬坡」可能只是多买了推理算力的幻觉。
把握程度:中等偏高。最强支撑是 Hamel 反复独立观察到"看 trace 是最有价值的部分",Camp A 内部无反对票。最弱环节仍是"私有 eval 要多大才算护城河"——它跟 ai-moat-2026 的"data moat 只在 mega-scale 成立"存在张力,且三个 Camp F 信源里有两个(Foody、Satya)是卖这条叙事的当事人,需要客户侧的独立证据才能锁死。
2026-08-24 再补一笔:六篇新料让我把「护城河在积累专有 trace 并转成内部 judge 的能力」这个判断,升级成一个更动态的版本——护城河可能不在攒出来的 eval set,而在 trace→eval→权重这条循环的转速(阵营 G 的核心主张,恰好也是"私有 eval 要攒到多大"这个老悬案的一种消解:如果资产是循环而不是存量,门槛问题就变成了速度问题)。同时 Camp F 的路线分裂让 Satya 测试的适用域缩小了:它只在「模型层保持商品化」的路线里有意义,对走 taste-in-weights 路线的公司无从施测。警惕项也要升级:本轮六篇说话人全是闭环卖方,「evals 是护城河」正在从论点变成行业销售话术——话术化不证伪它,但让下一份买方侧证据的权重变得更高。
还想知道什么
- 一个公司"手工 trace 标注 → 内部 eval-judge → 模型迭代"全闭环的真实案例:周期多长、人力多少、最终 quality 提升多少。Hamel 反复推这条 loop,但没有一篇访谈给出完整的事后 12 个月数据。
- "私有 eval 要攒到多大、多久"的定量门槛。 Satya 给了定性测试(换模型能否继续爬坡),但 Camp F 缺一个"多少条标注样本 / 多少个月 forward-deployed 之后,竞品复制成本才真正变高"的数字。这是 Camp F 对 a16z"data moat 只在 mega-scale 成立"那条最需要的反驳证据。
- Mercor TAM 演化数据:Foody 说市场受限于"人类能做、模型做不到的事的总量"——若模型在 2026 下半年吃掉一大块专家可证伪任务,Mercor 的 ARR 增速曲线会立刻给信号。
- SWE-Bench Pro / domain leaderboard / 私有 eval 三条路径的 head-to-head:同一个 agent 在三种 eval 体系下排名是否一致?不一致的话方差在哪?且按 Noam 的要求,比较必须固定在同一测试时算力预算下,否则全是噪声。
- 一个"我们靠 evals 在某 vertical 干掉某竞品"的客户侧第一人称叙述。 目前几乎所有"evals 是护城河"的论点都来自 vendor / 教学者(Mercor、Arize、Galileo、Hamel 课程、Microsoft,2026-08 后再加 Fireworks、Trajectory、LangChain)。缺一个买方视角的商战复盘。最接近的是 Ronak 转述的 Harvey/Nemotron 案例(法律指标全面提升、成本大降),但仍是训练供应商代述,不算数。
- 两条 own-intelligence 路线的对照实验。 同一家企业、同一批专有 trace:一条路走「私有 eval + 商品化模型层」(Satya 式,换模型自由),一条路走「蒸馏进自有权重」(Lin Qiao 式,5-10 倍推理成本优势但绑死权重),12 个月后比总成本、质量爬坡、供应商锁定。两派各有成套 talk track,语料里没有一个可比案例。
取材
核心 6 篇本轮按逐字稿全文重读、所有引用逐字稿核对:
- Mia Glaese & Olivia Watkins (OpenAI Frontier Evals — SWE-Bench-Dead) · 2026-02-26 ·
313ea6160e7181cfa590f64d2c2162af - Brendan Foody (Mercor — evals 劳动生意) · 2025-09-22 ·
276ea6160e7181edb819cd82fd44b5ac - Brendan Foody (Mercor — 私有 eval = 护城河 / 应用层无 defensibility) · 2026-06-06 ·
377ea6160e7181e38e78ea1048f812d1 - Satya Nadella (Microsoft — private eval = 最大 IP / Full-Stack Builder) · 2026-06-06 ·
377ea6160e718196bed4e563447eaa46 - Roman Chernin (Nebius — eval 是 cold-start 前置) · 2026-06-10 ·
37bea6160e71818187b7d7eaa67763b6 - Pratik Bhavsar (Galileo — Ranking Agentic LLMs) · 2025-07-25 ·
23bea6160e71810bb58fc297ea2443c4
沿用上一轮已逐字核对的 practitioner 引用(Camp A/D/E,原文均为第一人称逐字稿):
- Aman Khan (Arize) · 2025-08-25 ·
25aea6160e7181aaacb6e251289d1800 - Hamel Husain & Shreya Shankar · 2025-09-26 ·
27aea6160e7181bbb856ff899d12c8e4 - Hamel Husain (Crash Course, NurtureBoss) · 2025-10-04 ·
282ea6160e71813883fbee731f4d43e6 - Kyle Corbitt (OpenPipe — RULER) · 2025-10-17 ·
28fea6160e718186bfe4e90dde1d3a7d - Shunyu Yao & Harrison Chase (Language Agents) · 2025-07-25 ·
23bea6160e7181ff8d84f86195b392f6
2026-08-10 新增 3 篇(引用逐字核对):
- Noam Brown (OpenAI — benchmark grid 坏均衡 / 能力=算力预算的函数) · 2026-06 ·
38cea6160e71819db9edf75b0e6eb94d - Andy Beam & Rafa Gómez-Bombarelli (Lila Sciences — 自然/实验即终极 verifier) · 2026-07 ·
3a0ea6160e7181309f08d5135f750f56 - Nicole Brichtova & Oliver Wang (Google Nano Banana — 前沿实验室的 eyeball eval) · 2025-09 ·
276ea6160e7181d2b8baf486556602e3
同批另有 1 篇经逐字稿检查后未纳入:Baseten Philip Kiely & Ali Taha(推理工程;eval 仅以「量化保真度用 KL 散度而非 benchmark 分数衡量」边缘出现,3b2ea6160e71817bb609c9ceb8824f85)。
2026-08-24 新增 6 篇(全部纳入,引用逐字核对):
- Lin Qiao (Fireworks — taste-in-weights 反向路线 / vibe evaling→系统化 / reward=代码化 rubric) · 2026-08 ·
3c4ea6160e71815aa1a7dbcb3113e2ed - Brendan Foody (Mercor — RL 环境三件套 / 人类测量前沿之外 / 社交协作 eval 空白) · 2026-08 ·
3c4ea6160e718166bb67ff97d9a6f1b4 - Arjun Karanam (Trajectory — eval 取材生产流量 / 纠正行为信号) · 2026-08 ·
3c4ea6160e7181e0adc7cd856b471297 - Harrison Chase (LangChain — 转引 Satya 私有 eval 论 / trace flywheel 与 LangSmith Engine) · 2026-08 ·
3c4ea6160e71813ab82afaad1a81bade - Ronak Malde (Trajectory — 静态模型浪费纠正信号 / trace→evals 统一格式,swyx 访谈) · 2026-08 ·
3c4ea6160e718148b018d4952c6608e2 - Ronak Malde (Trajectory — trillion token problem / 标量 reward 信息量批判 / OPSD,大会演讲) · 2026-08 ·
3c4ea6160e7181d2ba5fef05c86fc589
注:阵营 F 的三篇(Foody-20VC、Satya、Nebius)当前未进入 topics.json 的 alias 成员表,是上一轮手工引入;建议后续 topics.py update 时把它们的 id 补进 interview_ids,以便 headless 全量重综能自动纳入。