复合系统胜过单一大模型:路由、编排与 Token 理性化 · Compound Systems Over the Monolith
主题综述
更新日志
- 2026-08-04 — 首次综述。基于 10 篇访谈:token maxing 阶段正在结束、"没有一个模型适合所有任务"已成跨阵营共识,真正的分歧在 harness 该绑定模型家族还是保持模型无关(Anthropic vs Factory/Cognition 正面相撞),以及路由到底是产品、生意还是会被模型内化的临时脚手架。
主流共识
一、Token maxing 的阶段性使命已经完成——现在进入"智能分配"阶段
语料里几乎每个人都在讲同一条时间线:先不计成本地推动采用,然后账单到期。Factory 的 Matan Grinberg 把这个演化讲成了四幕剧,并给出了最生动的一线细节——银行里的人拿 Opus 问天气:
"It's this token maxing where People are using like Opus for literally everything like what's the weather NSF Opus? Tell me I don't know like There are banks that we are working with Where they are spending literally hundreds of thousands of dollars a month on people asking things like, literally, what is the weather?"「这是代币最大化,在这里人们使用 Opus 几乎所有的事情,比如说天气是什么。告诉我,我不知道,比如说有些银行与我们合作,其中他们每月花费数十万美元让人问诸如天气是什么的问题。」Matan Grinberg (Factory) · Factory's Matan Grinberg: The Coming 'Dark Factory'
Merge 的 Gil Feig 描述了同一条曲线的财务终点——账单落到 CFO 桌上:
"You see everyone saying token-max, token-max, and that's really great in theory. It actually is great in practice, too. You're seeing a lot being built, but then the bill comes to the CFO, and it's actually really brutal and way worse than they expected."「你会看到每个人都在说令牌最大化,令牌最大化,这在理论上确实很好。实际上,这在实践中也很好。你会看到很多东西在构建,但账单却到了首席财务官那,实际上真的是非常残酷,比他们预期的要糟糕得多。」Gil Feig (Merge) · The Token-Maxxing Bill That Shocks Every CFO
Harvey 的 Gabe Pereyra 从垂直应用侧确认了拐点的时间:能力约束期结束、成本约束期开始,而且就发生在最近半年:
"Up until recently, we've always been capability constrained. And so we always wanted to use the largest model … But that has changed in the past six months where now we are consuming like a huge number of tokens. I think for some of the labs, we are like the largest consumer of embeddings."「直到最近,我们一直受到能力的限制。所以我们总是希望使用最大的模型……但在过去六个月中,这种情况发生了变化,现在我们正在消耗大量的令牌。我认为在某些实验室,我们是嵌入的最大消费者。」Gabe Pereyra (Harvey) · Harvey Co-Founder on the Token Pricing Reckoning
Anthropic 平台团队的 Angela Jiang 把"下一步"抽象成一句话——智能之后,优化维度只剩成本和速度:
"And as these models get more and more capable, you're going to hit levels of intelligence maxing that are there, that then you want to do the next dimension. And the next dimension after intelligence will either be cost or it will be speed."「随着这些模型变得越来越强大,你会达到智能极限层次。然后你想要寻找下一个维度。智慧之后的下一个维度将是成本或速度。」Angela Jiang (Anthropic) · Building an Ecosystem, not a Walled Garden
值得注意的是共识的边界:没有人主张"少用 AI"。Anthropic 的 Jiang 和 Merge、Factory 都强调不要用预算上限扼杀使用——要用更聪明的分配替代粗暴的封顶。Legora CTO Jacob Lauritzen 则点破了 token 排行榜这种管理动作的荒谬:
"having a leaderboard, a lot of people say this, get a leaderboard, bring up token users at performance reviews, and that leads to token maxing, which is people just burn tokens just to look good. That's a really stupid way to do anything."「设置一个排行榜,很多人说这个,得到排行榜,在绩效评审中提到代币使用者,这会导致代币最大化,让人们只是消耗代币来让自己看起来好。这是一种非常愚蠢的做法。」Jacob Lauritzen (Legora) · Inside Legora's Tech Stack
二、"没有一个模型适合所有任务"——复合系统成为一线产品的默认架构
这一条已经从观点变成了产品事实。Cognition 把 Devin 明确设计成复合模型系统,Scott Wu 的描述最完整:
"It turns out that there are different models that are good for different parts of these tasks, right? … Devin can use any of the different models it has in its arsenal, which include all of these models from Anthropic, OpenAI, Google, etc. But also our own models, right, or open source models out there. And it will, you know, dynamically go and choose these models for these tasks."「结果发现不同的模型适合完成这些任务的不同部分,对吧?……Devin 可以使用其武器库中不同的模型,包括来自 Anthropic 的所有模型,OpenAI、Google 等等。还有我们自己的模型,当然,或是其他开源模型。它会动态选择这些模型来完成这些任务。」Scott Wu (Cognition) · Scott Wu, Cognition
OpenAI 一侧(Kevin Weil)给出的内部实践几乎一模一样——编排模型指挥廉价专用小模型,而不是一个巨型 prompt 撞大运:
"You may have an initial model that's orchestrating and is like putting a plan together and understanding what you should do to answer the question. And then you have different models. Maybe some of them are cheaper models that are trained to do one thing really well. And the orchestration model is calling the other models and things. I don't see people doing that enough."「你可能有一个初始模型在协调,制定计划并理解你应该如何回答问题。然后你有不同的模型。也许其中一些是被训练来很好地执行一项工作的廉价模型。而协调模型正在调用其他模型和相关事宜。我不看到人们足够地这样做。」Kevin Weil (OpenAI) · AI Is Crossing the Frontier of Human Knowledge
Harvey 的基准测试(LAB)把它变成了可测量的结论,并给出关键性价比数字——前沿模型 3 倍价格只换 10–20% 性能:
"What we've seen is actually every different model is good at something different. And so with the initial results we saw, anthropics models are quite strong, but there's areas where 5.5 is better. There's some areas where open source is better. And increasingly, it's not just which model is the best, it's which model can solve the task at the lowest price point."「但我们看到的是确实每个不同的模型在某些方面都有其优势。因此,根据我们看到的初始结果,Anthropic 的模型相当强大,但在某些领域 5.5 更好。有些领域开源更好。而且越来越多的情况并不仅仅是哪个模型最好,而是哪个模型能以最低的价格解决任务。」Gabe Pereyra (Harvey) · Harvey Co-Founder on the Token Pricing Reckoning
" Opus 4.7 is three times more expensive than 5.5, but it's 10 or 20% more performant."「Opus 4.7 的价格是 5.5 的三倍,但其性能仅提高了 10% 或 20%。」Gabe Pereyra (Harvey) · 同上
三、开源模型 = frontier-minus-one,而这已经够用了(对很多任务)
Matan 给了最不含糊的表述和最硬的数字(Factory 路由器上开源 token 份额一年从 <1% 涨到两位数):
"everyone is comparing like GLM 5.2 to the latest model like Opus 4.8 or GPT 5.6. But really they should be compared to Opus 4.7 or GPT 5.5. Why? Because generally the open models come later and they're kind of a generation behind. … The question is, are the open models getting as good as like frontier minus one? And the answer is unequivocally yes"「每个人都在比较像 GLM 5.2 和最新的模型如 Opus 4.8 或 GPT 5.6。但实际上应该与 Opus 4.7 或 GPT 5.5 进行比较。为什么?因为通常开放模型来得较晚,它们有点滞后于一代。……问题是,开放模型的水平是否已达到像前沿模型减去一个的水平?答案是毫无疑问地是的」Matan Grinberg (Factory) · The Coming 'Dark Factory'
"At the beginning of the year, it was less than 1% of tokens went to open models. In the first quarter, it became a single-digit percent. It is now crossed into being a double-digit percent of tokens."「在年初,投入开放模型的代币不到 1%。在第一季度,这个比例变为个位数百分比。现在已经超过了代币的两位数百分比。」Matan Grinberg (Factory) · 同上
USV 的 Mike Mignano 从投资侧给出配套判断——企业里 80% 的非编码任务根本不需要前沿模型:
"Eighty percent of non-coding tasks in the enterprise can be done with models that are not at the frontier. I think if you're coding, you probably want to be leveraging frontier models."「企业中 80% 的非编码任务可以用并非前沿的模型完成。我认为如果你在编程,你可能想利用最前沿的模型。」Mike Mignano (USV) · Why Now is the Time for the Application Layer
分歧在哪
共识止步于"要用多个模型"。往下一层——harness 和模型该是什么关系、路由是不是一门生意、token maxing 该不该继续——阵营立刻分开。
分歧一 · Harness 该模型无关,还是绑定模型家族?(正面相撞)
这是本主题最硬的一条对立,双方都把话说得很满。Factory 的 Matan 主张多模型 harness 必然更强,还给出了机制(防过拟合)和实证(TerminalBench 上 Droid 一度跑赢 Claude Code/Codex):
"I think one thing that Naively, everyone believed initially was if you train the model and you build the harness, you're going to make them better together. And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better."「我认为一个简单的想法是,大家最初都相信如果你训练模型并构建系统,它们会一起变得更好。令我许多在 OpenAI 和 Anthropic 的朋友感到沮丧的是,这并不是真的。如果你构建一个支持不同模型的系统,那这个系统会表现更好。」Matan Grinberg (Factory) · The Coming 'Dark Factory'
"there's a sort of analog that emerges where it's what data is to a model, models are to a harness. Where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular."「由此出现了一种类比,数据对模型的意义,就像模型对系统的意义。当你让更多的模型接触到一个系统时,你会避免将该系统过拟合到特定模型的细节。」Matan Grinberg (Factory) · 同上
Cognition 的 Scott Wu 站同一边,且把中立本身当作商业模式("我们喜欢做瑞士"):
"Yeah, we like being Switzerland. Exactly. And so I think it's like an important thing of, you know, We are just as incentivized as they are to figure out how to make their token spend efficient, right? And so Devin is purposely meant to be a compound model system."「是的,我们喜欢做瑞士。正是这样。所以我认为这很重要,你知道,我们和他们一样有激励去找出如何让他们的代币花费变得高效,对吧?所以 Devin 是故意被设计成复合模型系统。」Scott Wu (Cognition) · Scott Wu, Cognition
Anthropic 的 Katelyn Lesse 的立场正好相反——harness 和 agentic 层就该跟模型家族协同调校,Vercel 的做法被她当作行业转向的证据:
"I think we have a strong belief that harnesses and just like the agentic layer should be tuned to the model family that you use it with. I think there was a period where people were kind of like, yeah, cool, I can like build a harness and build an agent and then just like plug in a different model underneath. And they were excited about routers from that perspective. And I think we started to see, like Verstel just did this with harness agent, for example, … come up a layer of abstraction and say, actually like plug in the whole harness and the whole agent that's tied to a model family"「我认为我们坚信,工具和代理层应该与您使用的模型系列进行调整。我认为曾经有一个时期,人们觉得,酷,我可以构建一个工具并构建一个代理,然后只需插入不同的模型。他们对此类路由器充满了兴奋。我认为我们开始看到,比如 Verstel 刚刚和 harness agent 一起做了这个……提出了一个抽象层,并表示,实际上可以将整个 harness 和与模型系列相关的整个 agent 插入进去」Katelyn Lesse (Anthropic) · Building an Ecosystem, not a Walled Garden
Angela Jiang 补上了平台边界——Anthropic 可以做路由,但只在 Claude 家族内部:
"I think the bit that we do feel really strongly about on the model routing front is like we are designing our platform for Claude and we want to make sure that Claude is great at like solving all these things. So we'll like restrict to that space rather than, you know, I don't think we're that interested in saying like, okay, and then, you know, you should route to a different model or whatever."「我们在模型路由方面非常坚定的观点是,我们正在为 Claude 设计我们的平台,我们希望确保 Claude 能很好地解决所有这些问题。所以我们将限制在该领域,而不是,你知道,我认为我们并不那么感兴趣于说,好吧,然后,你知道,你应该路由到不同的模型或其他什么。」Angela Jiang (Anthropic) · 同上
注意双方的利益位置恰好解释立场:模型无关性是 Factory/Cognition 对企业客户的核心卖点(Matan:"他们不想要任何人成为他们的单点故障"),而 lab 恰恰希望 harness 与模型"better together"——Matan 自己也承认这一点("从实验室的角度来看,你理想中希望它们能够更好地结合在一起,因为那样意味着你必须使用他们的安全带")。分歧是真实的,但它同时是一场卡位战。
分歧二 · Harness 本身是资产,还是会被模型吃掉的临时脚手架?
第三个阵营根本不接受上面那场争论的前提。AI2 的 Nathan Lambert 直接站"无 harness"派(与 Noam Brown 同队):
"I'm with Noam on no harnesses."
「我和 Noam 一样,不用 harness。」
"Yeah. I mean, harnesses are cool, but they're They're a handicap that's changing the learning dynamic substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses."「是的。我的意思是,harness 很酷,但它们是一种阻碍,会大大改变学习动态。所以这很好。好的演示,但我觉得核心推动力必须是没有 harness。」Nathan Lambert (AI2) · The RLVR Revolution
Merge 的 Gil Feig 从产品侧给出同构判断——自然语言即将成为可确定执行的编程语言,定制 harness/工作流构建器都不重要:
"You see a lot of people trying to build a custom harness or use something like a workflow builder to build agents that are repeatable and all of that. And I just ultimately think none of it matters because we're almost at the point now where English is the language that you use to tell an agent it will be deterministic very soon."「你会看到很多人尝试构建定制的工具或使用类似工作流构建器的东西来构建可重复的代理。我最终认为这些都没有意义,因为我们几乎到了一个点,即英语是你用来告诉代理的语言,很快就会确定性。」Gil Feig (Merge) · The Token-Maxxing Bill That Shocks Every CFO
有意思的是 Anthropic 的 Angela Jiang 部分承认了这个方向——引导型脚手架确实该删,但她的结论不是"无 harness",而是 harness 的重心上移到策略/协调层:
"If you look like two years ago, a lot of the harness was like a scaffold to kind of like Tell the model to go from point A to point B. … And now the models are actually very, very steerable. And so a lot of that steering, you can just put in the prompt, right? … if you have harnesses that are like designed to kind of do that kind of like steering, You can delete that part."「如果你回头看看两年前,很多的工具就像是一个支架,告诉模型从 A 点到 B 点。……现在模型实际上非常可控。所以很多的控制,你只需在提示中输入即可,对吧?……如果你有为了做这种控制而设计的工具,你可以删除那部分。」Angela Jiang (Anthropic) · Building an Ecosystem, not a Walled Garden
Kevin Weil 也在同一方向上留了口子——OpenAI 内部大量用模型集合,但"随着模型变强,对这种做法的需求会越来越少"。而 Harvey 站在完全相反的经验上:harness 是他们研究里最热的领域,因为垂直域的专业化和验证逻辑没法被通用模型吸收(Angela 自己也承认法律/金融的验证逻辑值得自持)。张力没有被调和:一边说 harness 是学习动态的枷锁、会被能力增长消化,一边说它是垂直价值的所在。
分歧三 · 路由是一门独立生意吗?
Mignano 看好路由层,还转述了一个奖励对齐的定价点子(选对模型才收费):
"I do think routing is interesting and important right now. … you're going to have companies that are singularly focused on this, companies like OpenRouter out of New York, which is doing some really interesting work."「我确实认为现在的路由是有趣且重要的。……所以你会看到专注于这一点的公司,如来自纽约的 OpenRouter,他们正在做一些非常有趣的工作。」Mike Mignano (USV) · Why Now is the Time for the Application Layer
主持人 Harry Stebbings 当场泼冷水:
"I think it's hard to see that $50 billion company built in routing alone, I have to say."「我认为只有通过路由层建立 500 亿美元的公司是很难的,我必须说。」Harry Stebbings · 同上
而做产品的人给出了第三种答案:路由不独立成业,而是内嵌在 harness/平台里的能力——Factory 的 router、Devin 的动态选模、Merge Gateway 的路由策略、Legora 眼里 Cursor 的存在理由("中立第三方帮你优化 token 支出")。Matan 甚至把它推演成一个市场机制的雏形("token 大炮"指向出价的供给方、按结果定价),但他自己承认这是前瞻推演,企业目前"只是从没有路由器到有路由器"的阶段。
分歧四 · 现在还该不该 token maxing?
共识说理性化,但 Mignano 给初创公司留了一个刺眼的例外——编码上继续把 token 支出拉满:
"But also, I think as a startup, you need every advantage you can get right now. And so if I, you know, if I were the CEO of a startup right now, I would actually still be pounding the table to maximize token spend on the right things, right? Definitely with coding"「但我也认为作为一家初创公司,你现在需要每一个优势。所以如果我是现在一家初创公司的首席执行官,我真的会继续强调在对的事情上最大化代币支出。对吧?绝对是针对编码的」Mike Mignano (USV) · Why Now is the Time for the Application Layer
Scott Wu 则反对把 token 消耗当成生产力指标本身(Cognition 卖的是项目级 ROI,不是 token);Legora 的规则是"性能优先、不看成本"但按机会成本算账。三种姿态(战略性拉满 / 绑定产出 / 性能优先)都自称理性化,说明"token 理性化"这个词已经在被各方按自己的商业模式重新定义。
都没说透的
- 路由器自己的判断成本没人算。 按任务复杂度路由,意味着要先花智能去评估复杂度——这层元判断用什么模型、错误路由(把难任务发给弱模型)的代价怎么计入,语料里没有任何人给出数字。Harvey 的"$20 → $20,000 同产品内成本方差"恰恰说明错判的代价可以是三个数量级。
- 质量信号从哪来。 复合系统的前提是知道"这个任务这个模型够用"——Harvey 用 LAB 自建了法律域的 ground truth,但通用场景下按性价比路由需要每任务的质量回归,Gabe 预言会长出一个围绕 token 账单的"审计/优化生态"(类比律所六分钟计费明细),可谁给 router 出题、谁审 router,没人接。
- Anthropic 的"策略层"叙事和市场行为的错位。 Lesse/Jiang 说低层 harness 没多少油水可榨、alpha 在 token 的元分配(advising vs executing vs dreaming、best-of-N 第三杠杆)——但整个市场(Factory、Cognition、Harvey、Cursor)仍在 harness 层激烈竞争且赚到了钱。要么策略层的 alpha 还没被验证,要么现有玩家都在错的层上竞争——两种读法差别巨大,没人正面对质。
- 开源追赶的资金可持续性。 "frontier-minus-one 且份额暴涨"被当作既成事实引用,但没人问:如果开源模型持续吃掉推理流量,下一代开源模型的训练成本由谁承担、激励是否可持续。
我的看法
判断(不是事实):复合系统是当下的正确工程答案,但它隐含一个可能不稳的前提——前沿模型之间的可替换性会持续走高。实验室的对抗动作已经开始(Anthropic 明说 harness 该绑定模型家族、路由只在 Claude 家族内做;Matan 转述实验室朋友对"multi-model harness 更强"的懊恼),如果某家实验室真的做出"harness 协同训练带来不可替代的性能差",模型无关派的地基会松动。我赌中期结果是分层稳态:路由/编排作为独立生意难以做大(Harry 的怀疑成立——OpenRouter 类会被夹在两头),但作为 Factory/Devin/Cursor/Merge 这类产品的内嵌能力是真护城河;而实验室会用"策略层 + 家族内路由"守住高毛利段。Token maxing → 理性化的叙事本身没有争议,真正的输赢在谁拥有那个做分配决策的位置。
把握程度:中等。"没有一个模型适合所有任务"有跨阵营实证(LAB 数据、Devin/Factory 的产品行为、OpenAI 内部实践)支撑,很扎实;但"路由难以独立成业"主要靠 Harry 一句质疑加结构推演,且 outcome-based 定价(Matan 的市场机制、Mignano 的 bounty 模型)如果真跑通,独立路由层的价值会被重估。
还想知道什么
- 各家 router 的质量回归数据。 Matan 给了开源 token 份额(<1% → 两位数),但没给"路由到便宜模型后任务成功率/返工率变化"——这是判断理性化是真省钱还是假省钱的关键一环。缺一篇讲路由错误代价的实证访谈。
- Anthropic strategies/meta-harness 的落地形态。 "给 token 分配 advising/dreaming/executing 角色"上线后,第三方还能在哪一层建 meta-harness——这决定分歧二的走向。
- harness-model 协同训练的硬证据。 有没有实验室能拿出"绑定 harness 比模型无关 harness 在同模型上高 X%"的可复现数字?Matan 给了反向轶事(TerminalBench),正向的还没人给。
- 一个真实的模型切换案例。 企业把主力模型从 A 家换到 B 家(或换到开源)的完整成本核算——迁移工程、评估重建、性能回归。模型无关叙事的含金量全在这个数字里。
取材
- Matan Grinberg (Factory) · 2026-07-27 ·
3aaea6160e7181ffb327c8ccc53f4069 - Katelyn Lesse & Angela Jiang (Anthropic) · 2026-07-17 ·
3a0ea6160e7181e38305eb41507be678 - Mike Mignano (USV) · 2026-07-06 ·
395ea6160e718195b011ea57525b44f7 - Kevin Weil (OpenAI) · 2026-06-30 ·
38fea6160e7181d9862cd6a455019494 - Scott Wu (Cognition) · 2026-06-30 ·
38fea6160e71818ea231e2e0899ae109 - Gabe Pereyra (Harvey) · 2026-06-22 ·
387ea6160e71816d928afff83182d03f - Jacob Lauritzen (Legora) · 2026-06-10 ·
37bea6160e7181df99b1d74b916bc9b5 - Shensi Ding & Gil Feig (Merge) · 2026-06-06 ·
377ea6160e7181118adde1b77a450870 - Nathan Lambert (AI2) · 2025-08-04 ·
245ea6160e7181b9aa4ddfcf99f76e04 - Will Brown (Prime Intellect) · 2025-06-18 ·
216ea6160e7181d68d72ede27715b1e1(弱相关:仅在思考/非思考模型路由的推测层面触及本主题)