复合系统胜过单一大模型:路由、编排与 Token 理性化 · Compound Systems Over the Monolith
主题综述
更新日志
- 2026-08-24 — 新增 3 篇(Dianne Penn/Anthropic、Sonya Huang/Sequoia、Arjun Karanam/Trajectory)。共识一被 Penn 复杂化:她从实验室产品侧正面接住"token maxing"一词并翻转——token 支出是输入、实验才是产出,被理性化的应是生产支出而非实验预算(同一判断也给分歧四添了第四种姿态);分歧一 Anthropic 阵营首次拿到内部实证叙事("frontier products for frontier models",Opus 4.5 与 Claude Code 互为放大器);共识三被 Huang 升级——开源不止 frontier-minus-one,自持栈 + 后训练 + 在线学习号称可在垂直域 better-than-frontier(她称这是 2026 年的新事实,归功 Kimi K3/GLM 5.2),该节标题相应改写;新开分歧五"编排租来的智能,还是拥有自己的权重"(sovereign AI 阵营:Not Your Weights, Not Your Product;路由器被降格为通往自持的中途站,这在结构上反而支持了 Harry 对独立路由生意的怀疑,已织入分歧三);分歧二注入 Trajectory 的第三种答案——harness 是持续学习系统里三个可写层之一,全局事实进权重、用户偏好留 context/harness,且"更新落到哪一层"的决策本身应被抽象掉。"都没说透的"新增 sovereign 栈总拥有成本、"better-than-frontier 快照 vs 移动目标"两条;"我的看法"下修了分层稳态的把握度并标注三个新声音的利益位置。三篇均纳入,无未纳入篇目。
- 2026-08-04 — 首次综述。基于 10 篇访谈:token maxing 阶段正在结束、"没有一个模型适合所有任务"已成跨阵营共识,真正的分歧在 harness 该绑定模型家族还是保持模型无关(Anthropic vs Factory/Cognition 正面相撞),以及路由到底是产品、生意还是会被模型内化的临时脚手架。
主流共识
一、Token maxing 的阶段性使命已经完成——现在进入"智能分配"阶段
语料里几乎每个人都在讲同一条时间线:先不计成本地推动采用,然后账单到期。Factory 的 Matan Grinberg 把这个演化讲成了四幕剧,并给出了最生动的一线细节——银行里的人拿 Opus 问天气:
"It's this token maxing where People are using like Opus for literally everything like what's the weather NSF Opus? Tell me I don't know like There are banks that we are working with Where they are spending literally hundreds of thousands of dollars a month on people asking things like, literally, what is the weather?"「这是代币最大化,在这里人们使用 Opus 几乎所有的事情,比如说天气是什么。告诉我,我不知道,比如说有些银行与我们合作,其中他们每月花费数十万美元让人问诸如天气是什么的问题。」Matan Grinberg (Factory) · Factory's Matan Grinberg: The Coming 'Dark Factory'
Merge 的 Gil Feig 描述了同一条曲线的财务终点——账单落到 CFO 桌上:
"You see everyone saying token-max, token-max, and that's really great in theory. It actually is great in practice, too. You're seeing a lot being built, but then the bill comes to the CFO, and it's actually really brutal and way worse than they expected."「你会看到每个人都在说令牌最大化,令牌最大化,这在理论上确实很好。实际上,这在实践中也很好。你会看到很多东西在构建,但账单却到了首席财务官那,实际上真的是非常残酷,比他们预期的要糟糕得多。」Gil Feig (Merge) · The Token-Maxxing Bill That Shocks Every CFO
Harvey 的 Gabe Pereyra 从垂直应用侧确认了拐点的时间:能力约束期结束、成本约束期开始,而且就发生在最近半年:
"Up until recently, we've always been capability constrained. And so we always wanted to use the largest model … But that has changed in the past six months where now we are consuming like a huge number of tokens. I think for some of the labs, we are like the largest consumer of embeddings."「直到最近,我们一直受到能力的限制。所以我们总是希望使用最大的模型……但在过去六个月中,这种情况发生了变化,现在我们正在消耗大量的令牌。我认为在某些实验室,我们是嵌入的最大消费者。」Gabe Pereyra (Harvey) · Harvey Co-Founder on the Token Pricing Reckoning
Anthropic 平台团队的 Angela Jiang 把"下一步"抽象成一句话——智能之后,优化维度只剩成本和速度:
"And as these models get more and more capable, you're going to hit levels of intelligence maxing that are there, that then you want to do the next dimension. And the next dimension after intelligence will either be cost or it will be speed."「随着这些模型变得越来越强大,你会达到智能极限层次。然后你想要寻找下一个维度。智慧之后的下一个维度将是成本或速度。」Angela Jiang (Anthropic) · Building an Ecosystem, not a Walled Garden
值得注意的是共识的边界:没有人主张"少用 AI"。Anthropic 的 Jiang 和 Merge、Factory 都强调不要用预算上限扼杀使用——要用更聪明的分配替代粗暴的封顶。Legora CTO Jacob Lauritzen 则点破了 token 排行榜这种管理动作的荒谬:
"having a leaderboard, a lot of people say this, get a leaderboard, bring up token users at performance reviews, and that leads to token maxing, which is people just burn tokens just to look good. That's a really stupid way to do anything."「设置一个排行榜,很多人说这个,得到排行榜,在绩效评审中提到代币使用者,这会导致代币最大化,让人们只是消耗代币来让自己看起来好。这是一种非常愚蠢的做法。」Jacob Lauritzen (Legora) · Inside Legora's Tech Stack
08-24 更新:这条时间线第一次被一位实验室产品负责人正面复杂化。 Anthropic 首位技术 PM Dianne Penn 被主持人拿 Gary Tan 的"现在每年烧 10 万美元 token 就是活在 2028 年"追问时,没有站队"拉满"或"理性化",而是把这个词本身拆了——token 支出只是输入,真正该设目标的是实验产出:
"I think I take more of like a almost product lens. It's almost like token spin is more the input. And really, the output is what you described of experimentation. And I think if we were orienting, like goals around experimentation, I feel like that, that might be the better framing of the outcomes."「我觉得我更偏向于一种几乎是产品的视角。这几乎是代币支出更多是输入。而实际输出就是你所描述的实验。我觉得如果我们围绕实验设定目标,我觉得那可能是结果的更好框架。」Dianne Penn (Anthropic) · Anthropic's first technical PM on token maxing
这与 Lauritzen 的排行榜批评是同一个洞察的两面:token 消耗本身不是目标函数——无论把它当 KPI 拉满还是当成本砍掉,都是在错的变量上做优化。Penn 的分工方案隐含着对共识一的修正:被理性化的应该是生产支出(银行问天气),而实验支出恰恰不该封顶,因为发现新兴能力(emergent capabilities 的不连续跃升)没有替代路径——"最有创造力的思想家、最优秀的原型创造者,确实花了很多时间与 Claude 合作"(逐字稿转述)。她还把 token 关注度写进了 PM 的手艺定义:
"Here, you have to sweat the tokens as much as you sweat the pixels."「在这里,你必须像对待像素一样认真对待令牌。」Dianne Penn (Anthropic) · 同上
二、"没有一个模型适合所有任务"——复合系统成为一线产品的默认架构
这一条已经从观点变成了产品事实。Cognition 把 Devin 明确设计成复合模型系统,Scott Wu 的描述最完整:
"It turns out that there are different models that are good for different parts of these tasks, right? … Devin can use any of the different models it has in its arsenal, which include all of these models from Anthropic, OpenAI, Google, etc. But also our own models, right, or open source models out there. And it will, you know, dynamically go and choose these models for these tasks."「结果发现不同的模型适合完成这些任务的不同部分,对吧?……Devin 可以使用其武器库中不同的模型,包括来自 Anthropic 的所有模型,OpenAI、Google 等等。还有我们自己的模型,当然,或是其他开源模型。它会动态选择这些模型来完成这些任务。」Scott Wu (Cognition) · Scott Wu, Cognition
OpenAI 一侧(Kevin Weil)给出的内部实践几乎一模一样——编排模型指挥廉价专用小模型,而不是一个巨型 prompt 撞大运:
"You may have an initial model that's orchestrating and is like putting a plan together and understanding what you should do to answer the question. And then you have different models. Maybe some of them are cheaper models that are trained to do one thing really well. And the orchestration model is calling the other models and things. I don't see people doing that enough."「你可能有一个初始模型在协调,制定计划并理解你应该如何回答问题。然后你有不同的模型。也许其中一些是被训练来很好地执行一项工作的廉价模型。而协调模型正在调用其他模型和相关事宜。我不看到人们足够地这样做。」Kevin Weil (OpenAI) · AI Is Crossing the Frontier of Human Knowledge
Harvey 的基准测试(LAB)把它变成了可测量的结论,并给出关键性价比数字——前沿模型 3 倍价格只换 10–20% 性能:
"What we've seen is actually every different model is good at something different. And so with the initial results we saw, anthropics models are quite strong, but there's areas where 5.5 is better. There's some areas where open source is better. And increasingly, it's not just which model is the best, it's which model can solve the task at the lowest price point."「但我们看到的是确实每个不同的模型在某些方面都有其优势。因此,根据我们看到的初始结果,Anthropic 的模型相当强大,但在某些领域 5.5 更好。有些领域开源更好。而且越来越多的情况并不仅仅是哪个模型最好,而是哪个模型能以最低的价格解决任务。」Gabe Pereyra (Harvey) · Harvey Co-Founder on the Token Pricing Reckoning
" Opus 4.7 is three times more expensive than 5.5, but it's 10 or 20% more performant."「Opus 4.7 的价格是 5.5 的三倍,但其性能仅提高了 10% 或 20%。」Gabe Pereyra (Harvey) · 同上
08-24 更新:Trajectory 的 Arjun Karanam 把"复合系统"的定义又往前推了一格——不只是多个模型的组合,而是权重、harness、context 三层构成的一个整体系统,而且是一个应该随使用而持续被优化的系统:
"But the way we view it is that The intelligence that your product is run off of is a system. It's a system that has many components and true continual learning is something that optimizes across that system and optimize it based on what parts of the system needs to be updated for the information that you're learning."「但我们认为,您产品所运行的智能是一个系统。它是一个有许多组成部分的系统,真正的持续学习是优化整个系统,并根据您学习的信息,确定系统中哪些部分需要更新。」Arjun Karanam (Trajectory) · Continual Learning: How AI Agents Get Better With Every Use
注意这句话悄悄改变了本主题的坐标轴:此前所有人讨论的复合系统是静态组合(选对模型、搭好 harness),Karanam 的版本是学习系统(每次交互决定哪一层该被更新)。如果这条路线跑通,"复合胜过单体"的理由会从性价比换轨到复利。
三、开源模型 = frontier-minus-one——而 sovereign 阵营已经喊出 better-than-frontier(垂直域)
Matan 给了最不含糊的表述和最硬的数字(Factory 路由器上开源 token 份额一年从 <1% 涨到两位数):
"everyone is comparing like GLM 5.2 to the latest model like Opus 4.8 or GPT 5.6. But really they should be compared to Opus 4.7 or GPT 5.5. Why? Because generally the open models come later and they're kind of a generation behind. … The question is, are the open models getting as good as like frontier minus one? And the answer is unequivocally yes"「每个人都在比较像 GLM 5.2 和最新的模型如 Opus 4.8 或 GPT 5.6。但实际上应该与 Opus 4.7 或 GPT 5.5 进行比较。为什么?因为通常开放模型来得较晚,它们有点滞后于一代。……问题是,开放模型的水平是否已达到像前沿模型减去一个的水平?答案是毫无疑问地是的」Matan Grinberg (Factory) · The Coming 'Dark Factory'
"At the beginning of the year, it was less than 1% of tokens went to open models. In the first quarter, it became a single-digit percent. It is now crossed into being a double-digit percent of tokens."「在年初,投入开放模型的代币不到 1%。在第一季度,这个比例变为个位数百分比。现在已经超过了代币的两位数百分比。」Matan Grinberg (Factory) · 同上
USV 的 Mike Mignano 从投资侧给出配套判断——企业里 80% 的非编码任务根本不需要前沿模型:
"Eighty percent of non-coding tasks in the enterprise can be done with models that are not at the frontier. I think if you're coding, you probably want to be leveraging frontier models."「企业中 80% 的非编码任务可以用并非前沿的模型完成。我认为如果你在编程,你可能想利用最前沿的模型。」Mike Mignano (USV) · Why Now is the Time for the Application Layer
08-24 更新:Sequoia 的 Sonya Huang 把这条线推过了一个质变点。 Matan 的主张是"frontier-minus-one 够用了"(够用 = 更便宜地达到可接受质量);Huang 在 Sequoia 的 sovereign AI 峰会上(台下 80 家 portfolio 公司创始人)宣布的是"自持栈可以更好"——不是省钱够用,而是垂直域性能反超前沿:
"And in large part, this is thanks to the newest open weight models, especially Kimi K3 and GLM 5.2 being extremely good. … So with strong post-training, prompt, harness engineering, online learning, you can actually reach better than frontier performance by owning your stack. And so this is new for 2026."「在很大程度上,要归功于最新的开放权重模型,尤其是 Kimi K3 和 GLM 5.2 特别优秀。……通过强大的后训练、提示、介质工程、在线学习,你可以通过拥有自己的堆栈达到比前沿更好的性能。所以这是 2026 年的新事物。」Sonya Huang (Sequoia) · How Companies Are Building Their Own Intelligence
两点保留:其一,这是动员大会上的论点,台上没有给任何基准数字(Harvey 的 LAB 是她引用的间接证据);其二,"better than frontier" 的前提是后训练 + harness 工程 + 在线学习的全套投入,与 Matan"直接路由到开源省钱"是两种成本结构。但方向性判断值得记录:开源叙事正在从"便宜的替代品"升级为"可塑的基材"——权重可得,所以可改,这是闭源 API 结构性给不了的(她的说法:闭源栈"更高的起点,但也是较低的上限")。
分歧在哪
共识止步于"要用多个模型"。往下一层——harness 和模型该是什么关系、路由是不是一门生意、token maxing 该不该继续、以及(08-24 新增)智能到底该租还是该拥有——阵营立刻分开。
分歧一 · Harness 该模型无关,还是绑定模型家族?(正面相撞)
这是本主题最硬的一条对立,双方都把话说得很满。Factory 的 Matan 主张多模型 harness 必然更强,还给出了机制(防过拟合)和实证(TerminalBench 上 Droid 一度跑赢 Claude Code/Codex):
"I think one thing that Naively, everyone believed initially was if you train the model and you build the harness, you're going to make them better together. And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better."「我认为一个简单的想法是,大家最初都相信如果你训练模型并构建系统,它们会一起变得更好。令我许多在 OpenAI 和 Anthropic 的朋友感到沮丧的是,这并不是真的。如果你构建一个支持不同模型的系统,那这个系统会表现更好。」Matan Grinberg (Factory) · The Coming 'Dark Factory'
"there's a sort of analog that emerges where it's what data is to a model, models are to a harness. Where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular."「由此出现了一种类比,数据对模型的意义,就像模型对系统的意义。当你让更多的模型接触到一个系统时,你会避免将该系统过拟合到特定模型的细节。」Matan Grinberg (Factory) · 同上
Cognition 的 Scott Wu 站同一边,且把中立本身当作商业模式("我们喜欢做瑞士"):
"Yeah, we like being Switzerland. Exactly. And so I think it's like an important thing of, you know, We are just as incentivized as they are to figure out how to make their token spend efficient, right? And so Devin is purposely meant to be a compound model system."「是的,我们喜欢做瑞士。正是这样。所以我认为这很重要,你知道,我们和他们一样有激励去找出如何让他们的代币花费变得高效,对吧?所以 Devin 是故意被设计成复合模型系统。」Scott Wu (Cognition) · Scott Wu, Cognition
Anthropic 的 Katelyn Lesse 的立场正好相反——harness 和 agentic 层就该跟模型家族协同调校,Vercel 的做法被她当作行业转向的证据:
"I think we have a strong belief that harnesses and just like the agentic layer should be tuned to the model family that you use it with. I think there was a period where people were kind of like, yeah, cool, I can like build a harness and build an agent and then just like plug in a different model underneath. And they were excited about routers from that perspective. And I think we started to see, like Verstel just did this with harness agent, for example, … come up a layer of abstraction and say, actually like plug in the whole harness and the whole agent that's tied to a model family"「我认为我们坚信,工具和代理层应该与您使用的模型系列进行调整。我认为曾经有一个时期,人们觉得,酷,我可以构建一个工具并构建一个代理,然后只需插入不同的模型。他们对此类路由器充满了兴奋。我认为我们开始看到,比如 Verstel 刚刚和 harness agent 一起做了这个……提出了一个抽象层,并表示,实际上可以将整个 harness 和与模型系列相关的整个 agent 插入进去」Katelyn Lesse (Anthropic) · Building an Ecosystem, not a Walled Garden
Angela Jiang 补上了平台边界——Anthropic 可以做路由,但只在 Claude 家族内部:
"I think the bit that we do feel really strongly about on the model routing front is like we are designing our platform for Claude and we want to make sure that Claude is great at like solving all these things. So we'll like restrict to that space rather than, you know, I don't think we're that interested in saying like, okay, and then, you know, you should route to a different model or whatever."「我们在模型路由方面非常坚定的观点是,我们正在为 Claude 设计我们的平台,我们希望确保 Claude 能很好地解决所有这些问题。所以我们将限制在该领域,而不是,你知道,我认为我们并不那么感兴趣于说,好吧,然后,你知道,你应该路由到不同的模型或其他什么。」Angela Jiang (Anthropic) · 同上
08-24 更新:Anthropic 阵营第一次拿出了内部实证叙事。 此前"better together"主要是 Lesse 的立场宣示(外加 Vercel 转向作旁证),Dianne Penn 给出了实验室视角下最具体的案例——Opus 4.5 与 Claude Code 互为成立条件:
"One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models. … I think Opus 45 wouldn't have had that moment without a product like Claude Code. And Claude Code, I think, wouldn't have had that type of adoption accelerated without Opus 45."「我们团队常说,你需要前沿产品才能拥有前沿模型,让人们感受到前沿模型的魔力。……我认为没有像 Claude Code 这样的产品,Opus 4.5 不会有那个时刻。Claude Code,我认为,没有 Opus 4.5 的话,它的采用速度不会那么快。」Dianne Penn (Anthropic) · Anthropic's first technical PM on token maxing
严格说这仍是相关性叙事而非可复现数字("还想知道什么"里要的硬证据依然缺席),但它把 Anthropic 阵营的论点从"应然"(harness 该绑定家族)推进到了"实然"(我们的旗舰产品和旗舰模型确实是互相成就的)。注意双方的利益位置恰好解释立场:模型无关性是 Factory/Cognition 对企业客户的核心卖点(Matan:"他们不想要任何人成为他们的单点故障"),而 lab 恰恰希望 harness 与模型"better together"——Matan 自己也承认这一点("从实验室的角度来看,你理想中希望它们能够更好地结合在一起,因为那样意味着你必须使用他们的安全带")。分歧是真实的,但它同时是一场卡位战——而且(见分歧五)现在有第三方干脆拒绝了"模型是租来的"这个双方共享的前提。
分歧二 · Harness 本身是资产,还是会被模型吃掉的临时脚手架?
第三个阵营根本不接受上面那场争论的前提。AI2 的 Nathan Lambert 直接站"无 harness"派(与 Noam Brown 同队):
"I'm with Noam on no harnesses."
「我和 Noam 一样,不用 harness。」
"Yeah. I mean, harnesses are cool, but they're They're a handicap that's changing the learning dynamic substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses."「是的。我的意思是,harness 很酷,但它们是一种阻碍,会大大改变学习动态。所以这很好。好的演示,但我觉得核心推动力必须是没有 harness。」Nathan Lambert (AI2) · The RLVR Revolution
Merge 的 Gil Feig 从产品侧给出同构判断——自然语言即将成为可确定执行的编程语言,定制 harness/工作流构建器都不重要:
"You see a lot of people trying to build a custom harness or use something like a workflow builder to build agents that are repeatable and all of that. And I just ultimately think none of it matters because we're almost at the point now where English is the language that you use to tell an agent it will be deterministic very soon."「你会看到很多人尝试构建定制的工具或使用类似工作流构建器的东西来构建可重复的代理。我最终认为这些都没有意义,因为我们几乎到了一个点,即英语是你用来告诉代理的语言,很快就会确定性。」Gil Feig (Merge) · The Token-Maxxing Bill That Shocks Every CFO
有意思的是 Anthropic 的 Angela Jiang 部分承认了这个方向——引导型脚手架确实该删,但她的结论不是"无 harness",而是 harness 的重心上移到策略/协调层:
"If you look like two years ago, a lot of the harness was like a scaffold to kind of like Tell the model to go from point A to point B. … And now the models are actually very, very steerable. And so a lot of that steering, you can just put in the prompt, right? … if you have harnesses that are like designed to kind of do that kind of like steering, You can delete that part."「如果你回头看看两年前,很多的工具就像是一个支架,告诉模型从 A 点到 B 点。……现在模型实际上非常可控。所以很多的控制,你只需在提示中输入即可,对吧?……如果你有为了做这种控制而设计的工具,你可以删除那部分。」Angela Jiang (Anthropic) · Building an Ecosystem, not a Walled Garden
Kevin Weil 也在同一方向上留了口子——OpenAI 内部大量用模型集合,但"随着模型变强,对这种做法的需求会越来越少"。而 Harvey 站在完全相反的经验上:harness 是他们研究里最热的领域,因为垂直域的专业化和验证逻辑没法被通用模型吸收(Angela 自己也承认法律/金融的验证逻辑值得自持)。张力没有被调和:一边说 harness 是学习动态的枷锁、会被能力增长消化,一边说它是垂直价值的所在。
08-24 更新:Trajectory 的 Arjun Karanam 给出了第三种答案,两边各否定一半。 他不同意"harness 会被模型吃掉"——因为有一类知识本来就不该进权重;也不同意"harness 是资产本身"——因为 harness 只是学习系统的三个可写层之一。他的划分标准是反馈的适用范围:全局为真的教训写进权重,局部偏好留在 context/harness:
"There's some information that's probably globally accurate, globally true. Things like a tool call repeatedly failing when trying to call a certain tool. That's probably relevant to everybody. And so this should be trained into the model and the model should get better at learning this tool. … A certain user is like, I never want to use a sub-agent. Please, please, please don't do it. That's probably not something you should train to the model and leave to the context."「有些信息可能是全球准确、全球真实的。比如某个工具在尝试调用特定工具时反复失败。这可能与每个人都相关。所以这应该被训练进模型,模型应该在学习这个工具时变得更好。……某个用户会说,我从来不想使用子代理。请、请、请不要这样做。这可能不是你应该训练到模型中的内容,留待上下文处理。」Arjun Karanam (Trajectory) · Continual Learning: How AI Agents Get Better With Every Use
更激进的是下一步:他认为"更新落到哪一层"这个决策本身不该由人做,应该像内存管理一样被抽象掉——
"We no longer think of like, where in your RAM or where in your hard disk to save. That's the level that's abstracted based on what makes the most sense. We think about models versus harnesses versus context in the same way. … This is a scientific problem that can be solved. Let's do that. And then let's abstract that away."「我们不再考虑在 RAM 或硬盘上保存的位置。这是一个根据最合理的数据进行抽象的层次。我们以同样的方式考虑模型、工具和上下文。……这是一个可以解决的科学问题。让我们这样做。然后将其抽象。」Arjun Karanam (Trajectory) · 同上
这个立场与 Jiang 的"删掉引导型脚手架"兼容(他也说要从"限制流程"转向"编排产品原语、让 agent 自己发挥"),与 Lambert 的"无 harness"直接冲突(Lambert 要的恰是让权重吞掉一切学习信号,Karanam 明确保留 harness 作为局部知识的居所——并援引 Harvey 的 per-org/per-customer 分层作为同路证据)。如果说分歧二原来的问题是"harness 会不会死",Trajectory 把问题改写成了"harness 是缓存层还是持久层"——它不死,但也不再是护城河本体,护城河挪到了那个决定写哪一层的学习闭环上。
分歧三 · 路由是一门独立生意吗?
Mignano 看好路由层,还转述了一个奖励对齐的定价点子(选对模型才收费):
"I do think routing is interesting and important right now. … you're going to have companies that are singularly focused on this, companies like OpenRouter out of New York, which is doing some really interesting work."「我确实认为现在的路由是有趣且重要的。……所以你会看到专注于这一点的公司,如来自纽约的 OpenRouter,他们正在做一些非常有趣的工作。」Mike Mignano (USV) · Why Now is the Time for the Application Layer
主持人 Harry Stebbings 当场泼冷水:
"I think it's hard to see that $50 billion company built in routing alone, I have to say."「我认为只有通过路由层建立 500 亿美元的公司是很难的,我必须说。」Harry Stebbings · 同上
而做产品的人给出了第三种答案:路由不独立成业,而是内嵌在 harness/平台里的能力——Factory 的 router、Devin 的动态选模、Merge Gateway 的路由策略、Legora 眼里 Cursor 的存在理由("中立第三方帮你优化 token 支出")。Matan 甚至把它推演成一个市场机制的雏形("token 大炮"指向出价的供给方、按结果定价),但他自己承认这是前瞻推演,企业目前"只是从没有路由器到有路由器"的阶段。
08-24 更新:sovereign 阵营给了路由第四种定位——中途站。 Sonya Huang 的技术路线图把 router 放在 evals 之后、post-training 之前,是自持旅程的第二站而不是终点:
"Next, we see companies starting to play with model routers, with harnesses. Some companies find they can get good performance with out-of-the-box models. Others are finding strong performance gains from post-training, in some rarer cases needing to move into mid-training, pre-training. And then finally, setting that machine up so that live customer data is actually creating a feedback loop where your model, your intelligence is improving with every customer interaction."「接下来,我们看到公司开始玩模型路由器和介质。一些公司发现,他们可以通过开箱即用的模型获得良好的性能。其他公司则发现,从后训练中获得了强大的性能提升,在一些较少见的情况下,甚至需要转向中间训练、前训练。然后最后,设置该机器,使得实时客户数据实际上创造了一个反馈回路,在每次客户互动中都能提升你的模型、你的智能。」Sonya Huang (Sequoia) · How Companies Are Building Their Own Intelligence
Trajectory 的 Karanam 同场呼应(并转述 Gabe Pereyra 同一观点):路由器会"在将智能路由到任务的确切能力中发挥非常重要的作用",他的四个愿望清单里明确包含"尝试使用路由器"(逐字稿有此句)。注意这两票都是看多路由、看空路由生意:路由重要,但它是每家公司学习闭环里的一个执行器,且最先进的客户会继续往 post-training/在线学习走。这在结构上支持 Harry 的怀疑——如果 router 只是旅程的第二站,独立路由公司的客户会不断"毕业"离开。
分歧四 · 现在还该不该 token maxing?
共识说理性化,但 Mignano 给初创公司留了一个刺眼的例外——编码上继续把 token 支出拉满:
"But also, I think as a startup, you need every advantage you can get right now. And so if I, you know, if I were the CEO of a startup right now, I would actually still be pounding the table to maximize token spend on the right things, right? Definitely with coding"「但我也认为作为一家初创公司,你现在需要每一个优势。所以如果我是现在一家初创公司的首席执行官,我真的会继续强调在对的事情上最大化代币支出。对吧?绝对是针对编码的」Mike Mignano (USV) · Why Now is the Time for the Application Layer
Scott Wu 则反对把 token 消耗当成生产力指标本身(Cognition 卖的是项目级 ROI,不是 token);Legora 的规则是"性能优先、不看成本"但按机会成本算账。08-24 起是四种姿态:Dianne Penn 补上了第四种——token 支出是实验的输入而非目标,该设 KPI 的是实验产出(见共识一的引语)。方向与 Mignano 的"战略性拉满"接近,但理由不同:不是"初创公司需要每一个优势"的竞争逻辑,而是"新兴能力的发现没有替代路径"的研发逻辑。四种姿态都自称理性化,说明"token 理性化"这个词已经在被各方按自己的商业模式重新定义。
分歧五 · 编排租来的智能,还是拥有自己的权重?(新阵营进场)
分歧一的双方共享一个前提:模型是从实验室租来的,争的只是 harness 跟它绑多紧。Sequoia 的 Sonya Huang 代表的 sovereign AI 阵营拒绝这个前提——复合系统的终局不是聪明地租,而是把关键部分的权重拿到自己手里:
"I hereby present the AI version of this meme, Not Your Weights, Not Your Product. I think that for a product to be truly yours, I think it's reasonable to think that you need to be able to control and custody your own weights."「我在这里呈现这个梗的 AI 版本:没有你的权重,就没有你的产品。我认为,要使产品真正属于你,合理的想法是你需要能够控制和保管自己的权重。」Sonya Huang (Sequoia) · How Companies Are Building Their Own Intelligence
她的战场判断把本主题里 Harvey/Factory 的角色重新命名了:应用公司就是新一代实验室("最热的新实验室实际上是来自像哈维这样的公司的应用研究,像工厂、Glean……",逐字稿有此句),竞争从应用层下沉到智能层本身——恰在她演讲前一天,Harvey 宣布成立 Harvey Research。但要防止把这个阵营读过头,她自己划了边界——sovereign 是光谱不是二元,前沿 agentic 任务照旧租:
"An important nuance here is we are definitely not telling our companies to get off Opus or GPT. That is definitely not the message. For coding agents, for desktop work, for frontier level APIs, the closed model APIs are wonderful."「一个重要的细微差别是,我们并不是在告诉我们的公司要放弃 Opus 或 GPT。这绝对不是我们的意思。对于编码代理、桌面工作和前沿级的 API,封闭模型的 API 是很棒的选择。」Sonya Huang (Sequoia) · 同上
她给的 own-vs-rent 判据是四个因子:成本(COGS 占比)、延迟(是不是 P0)、性能(自有数据调优能否反超)、专有数据浓度——tab autocomplete 已经整体迁到自持模型,agent 仍然整体租用。这个切分与 Mignano 的"80% 非编码任务不需要前沿"惊人地互补:高频、低延迟、域数据浓的段位归 sovereign,前沿推理段位归实验室 API。Trajectory 的 Karanam 则把"为什么要自持"从战略论证换成了技术论证——不自持权重,就没法做持续学习:
"So start getting comfortable running on open weights because that is what unlocks the door to owning your weights and then continually improving on top of them."「所以开始习惯使用开放权重,因为这将打开拥有您权重的大门,并随后在其上不断改进。」Arjun Karanam (Trajectory) · Continual Learning: How AI Agents Get Better With Every Use
利益位置照例要标注:这两票来自同一个讲台——Sequoia 在动员自己的 portfolio 走一条资本开支更重的路(对 venture 回报有利),Trajectory 卖的就是持续学习平台(Karanam 是 Sequoia 峰会的压轴讲者)。Huang 也坦承自持是"打开了潘多拉的盒子"(干净的 API 调用变成选基座、后训练、配 harness、建数据引擎的全套工程)。但阵营是真实的:它给分歧一提供了第三个答案(harness 和模型都该是你自己的——"生产堆栈是模型之上的一个工具",两层一起自持),给分歧二提供了新的赌注结构(harness 的价值不在于它本身,而在于它是自持栈里你能完全控制的那一层),也解释了共识三里开源份额暴涨的需求端动力。
都没说透的
- 路由器自己的判断成本没人算。 按任务复杂度路由,意味着要先花智能去评估复杂度——这层元判断用什么模型、错误路由(把难任务发给弱模型)的代价怎么计入,语料里没有任何人给出数字。Harvey 的"$20 → $20,000 同产品内成本方差"恰恰说明错判的代价可以是三个数量级。(08-24 补:Karanam 的"把更新落层决策抽象掉"是对这层元判断的正面宣战——但目前是研究愿景,不是已交付的数字。)
- 质量信号从哪来。 复合系统的前提是知道"这个任务这个模型够用"——Harvey 用 LAB 自建了法律域的 ground truth,但通用场景下按性价比路由需要每任务的质量回归,Gabe 预言会长出一个围绕 token 账单的"审计/优化生态"(类比律所六分钟计费明细),可谁给 router 出题、谁审 router,没人接。(08-24 补:Trajectory 给了半个答案——eval 从生产流量抽样、捕捉 undo/edit/retry 等纠正信号而非点赞点踩;但这套办法预设你拥有产品和栈,租用多模型场景下的第三方审计依然无人认领。)
- Anthropic 的"策略层"叙事和市场行为的错位。 Lesse/Jiang 说低层 harness 没多少油水可榨、alpha 在 token 的元分配(advising vs executing vs dreaming、best-of-N 第三杠杆)——但整个市场(Factory、Cognition、Harvey、Cursor)仍在 harness 层激烈竞争且赚到了钱。要么策略层的 alpha 还没被验证,要么现有玩家都在错的层上竞争——两种读法差别巨大,没人正面对质。
- 开源追赶的资金可持续性。 "frontier-minus-one 且份额暴涨"被当作既成事实引用,但没人问:如果开源模型持续吃掉推理流量,下一代开源模型的训练成本由谁承担、激励是否可持续。(08-24 补:sovereign 阵营让这个问题更尖锐——Huang 的整套论证建立在 Kimi K3/GLM 5.2"特别优秀"之上,等于把 80 家公司的技术路线押在别人继续放权重的善意上,台上无人提及。)
- Sovereign 栈的总拥有成本没人报账。(08-24 新增)Huang 的四驱动力里"成本"排第一,但潘多拉盒子打开后的另一侧——研究团队、数据引擎、GPU、eval 维护——没人把它和被替代的 API 账单放进同一张表。唯一的数据点是"Harvey 只有七人研究团队",而那恰是幸存者展台。
- "Better than frontier" 是对移动目标的快照。(08-24 新增)前沿每几个月跳一级,开源基座按 Matan 的口径固定滞后一代;自持模型在每次前沿跃迁后要花多少钱、多少周期重新反超,Huang 的路线图里没有这一项。如果答案是"每次都要全量重来",sovereign 的性能论证只在两次前沿发布的间隙里成立。
我的看法
判断(不是事实):复合系统是当下的正确工程答案,但它隐含一个可能不稳的前提——前沿模型之间的可替换性会持续走高。实验室的对抗动作已经开始(Anthropic 明说 harness 该绑定模型家族、路由只在 Claude 家族内做;Dianne Penn 补上了 Opus 4.5 × Claude Code 互为成立条件的内部叙事;Matan 转述实验室朋友对"multi-model harness 更强"的懊恼),如果某家实验室真的做出"harness 协同训练带来不可替代的性能差",模型无关派的地基会松动。我原本赌的中期结果是分层稳态:路由/编排作为独立生意难以做大(Harry 的怀疑成立——OpenRouter 类会被夹在两头),但作为 Factory/Devin/Cursor/Merge 这类产品的内嵌能力是真护城河;而实验室会用"策略层 + 家族内路由"守住高毛利段。
08-24 修订:分层稳态的框架保留,但"谁拥有分配决策的位置"从两极变成了三极——实验室(家族内路由 + 策略层)、编排厂商(模型无关 harness)、以及现在正式进场的应用公司自己(sovereign 栈 + 在线学习)。边界线我现在赌这样划:高频、低延迟、域数据浓的任务段被 sovereign 栈吃掉(tab autocomplete 已是既成事实),前沿 agentic 段继续租(Huang 自己都这么劝);真正的新变量是持续学习——如果"系统随使用复利"跑通,护城河会从"选对模型"整体迁移到"拥有学习闭环",届时路由降格为闭环里的一个执行器,分歧一那场 harness 卡位战的赌注也会缩水。但这一层判断的证据基础明显更弱:三个新声音里两个是在 Sequoia 的动员大会上带货(Sequoia 押 portfolio 走重资产路线、Trajectory 卖持续学习平台),"better than frontier"台上没有给出任何基准数字,Dianne 的 better-together 也仍是内部相关性叙事。
把握程度:原有部分中等——"没有一个模型适合所有任务"有跨阵营实证(LAB 数据、Devin/Factory 的产品行为、OpenAI 内部实践)支撑,很扎实;"路由难以独立成业"现在多了一层结构支撑(sovereign 路线图把 router 定位为客户会毕业离开的中途站),略有上调。新增的 sovereign/持续学习判断中等偏低——方向可信(四因子切分逻辑自洽、与 Mignano 的 80% 判断互补),量级存疑(无数字、重利益、快照对移动目标)。
还想知道什么
- 各家 router 的质量回归数据。 Matan 给了开源 token 份额(<1% → 两位数),但没给"路由到便宜模型后任务成功率/返工率变化"——这是判断理性化是真省钱还是假省钱的关键一环。缺一篇讲路由错误代价的实证访谈。
- Anthropic strategies/meta-harness 的落地形态。 "给 token 分配 advising/dreaming/executing 角色"上线后,第三方还能在哪一层建 meta-harness——这决定分歧二的走向。
- harness-model 协同训练的硬证据。 有没有实验室能拿出"绑定 harness 比模型无关 harness 在同模型上高 X%"的可复现数字?Matan 给了反向轶事(TerminalBench),Dianne Penn 给了正向案例叙事(Opus 4.5 × Claude Code),但仍然没人给数字。
- 一个真实的模型切换案例。 企业把主力模型从 A 家换到 B 家(或换到开源)的完整成本核算——迁移工程、评估重建、性能回归。模型无关叙事的含金量全在这个数字里。
- 一个跨两代前沿发布仍保持域内领先的自持模型。(08-24 新增)"better than frontier"要成立为趋势而非快照,需要至少一个案例:自持栈在前沿模型跳级后 N 个月内以可承受成本重新反超。这直接检验分歧五的成色。
- 持续学习的更新分布实证。(08-24 新增)Trajectory 主张全局事实进权重、局部偏好进 context/harness——真实生产系统里这两类更新的比例是多少?如果 95% 的学习信号都落在 context/harness 层,那"必须自持权重才能持续学习"的论证会被大幅削弱。
取材
- Sonya Huang (Sequoia) · 2026-08-22 ·
3c4ea6160e7181e3adf3cb7c4a57df0b - Arjun Karanam (Trajectory) · 2026-08-22 ·
3c4ea6160e7181e0adc7cd856b471297 - Dianne Penn (Anthropic) · 2026-08-04 ·
3b2ea6160e7181c39fa3f0b4340ee088 - Matan Grinberg (Factory) · 2026-07-27 ·
3aaea6160e7181ffb327c8ccc53f4069 - Katelyn Lesse & Angela Jiang (Anthropic) · 2026-07-17 ·
3a0ea6160e7181e38305eb41507be678 - Mike Mignano (USV) · 2026-07-06 ·
395ea6160e718195b011ea57525b44f7 - Kevin Weil (OpenAI) · 2026-06-30 ·
38fea6160e7181d9862cd6a455019494 - Scott Wu (Cognition) · 2026-06-30 ·
38fea6160e71818ea231e2e0899ae109 - Gabe Pereyra (Harvey) · 2026-06-22 ·
387ea6160e71816d928afff83182d03f - Jacob Lauritzen (Legora) · 2026-06-10 ·
37bea6160e7181df99b1d74b916bc9b5 - Shensi Ding & Gil Feig (Merge) · 2026-06-06 ·
377ea6160e7181118adde1b77a450870 - Nathan Lambert (AI2) · 2025-08-04 ·
245ea6160e7181b9aa4ddfcf99f76e04 - Will Brown (Prime Intellect) · 2025-06-18 ·
216ea6160e7181d68d72ede27715b1e1(弱相关:仅在思考/非思考模型路由的推测层面触及本主题)