主题综述

Vibe Coding:真范式还是新工具 · Vibe Coding Shift

主题综述

更新日志

主流共识

工具已经从根本上改变了工程师的工作单元。这一点没有反对票。

据访谈中引用的一份 semi-analysis 报告:Claude Code 已贡献 GitHub 上约 4% 的公开提交,该报告还预测到 2026 年底这一比例会涨到约五分之一。(注:这是主持人引用的第三方报告数据,非 Boris 本人原话。)

"95% of engineers use Codex. 100% of our PRs are reviewed by Codex daily as well. … Engineers who tend to use Codex more open way more PRs. So they're actually opening 70% more PRs than the engineers who aren't using Codex as much. And the gap is widening."
「95% 的工程师在用 Codex。我们 100% 的 PR 每天也都由 Codex 审查。……更常用 Codex 的工程师会开多得多的 PR。所以他们开的 PR 实际上比不那么常用 Codex 的工程师多 70%。而且差距在扩大。」

2026 年年中,这个共识从判断变成了计量。供给侧:Fiona Fung 那期开场,主持人 Lenny 引用 Anthropic 前一天发布的官方推文——Anthropic 工程师人均每季度 ship 的代码量已是 2025 年的 8 倍(注:数据出自 Anthropic 官方推文、由主持人转述;Fiona 本人的评论见下文阵营 A)。买方侧,Benchmark 的 Ev Randle 给出了市场愿意为这个位移付的价:

"We have developers that are spending $3,000 per month themselves, each on Cloud Code. It's like, wow, okay, so that's $36,000 per developer."
「我们有开发者每月在 Cloud Code 上花费 3,000 美元。哇,这样一来,每个开发者的花费就是 36,000 美元。」
Ev Randle · Benchmark's AI Bets

"人均/总量产出数倍化"至此已有两家公司独立报数:Anthropic(人均 8x/季,官方推文)、OpenAI(Codex 重度用户 +70% PR)。

2026-08-24 边界注记)这个计量共识第一次被从内部划出了职能边界:报数的全部是工程职能。OpenAI 产品设计负责人 Ian Silber 被问到设计侧时自报——设计并没有同步数倍化:

"Engineers have 10 or sometimes like 100X their productivity, but our design team hasn't because the design process still takes time. It's still sometimes you have to try a bunch of things and throw them out. Still very messy and fluid."
「工程师的生产力提高了 10 倍,有时甚至是 100 倍,但我们的设计团队并没有,因为设计过程仍然需要时间。有时你仍然需要尝试很多事情,然后再淘汰它们。仍然是非常混乱和流动的。」

"生成快了"只在产出物可由模型直接生成的职能上兑现为倍数;设计的瓶颈在试错与判断回路,本来就不在打字上。这不推翻共识,但给"人均产出数倍化"划了适用范围——它目前是工程侧数据,不是全职能事实。

第二点共识:工程师的角色在朝"管理代理"方向位移——即使各家对位移到什么终点有不同想象,"engineer-as-tech-lead-of-agents"是普遍的近期描述。本次更新后,这个位移有了两个更具体的一线形态:Fiona Fung 描述 Anthropic 内部正整体转向 async(routines 在夜里替她生成 prompt、跑 agent,人早上起来只审 PR),Kevin Weil 则把"个人多线程"当作新的基本功(见阵营 B)。

第三点共识(2026-07 新增):瓶颈从"生成"位移到了"验证"。这原本更接近阵营 B(Embiricos)的独家论点,现在阵营 A 的大本营也这么说了——Fiona Fung 的原话:

"coding is no longer the bottleneck. … Now, not only engineers, we also have designers, PMs, everybody on the Claude Code team checks in code. … but also the throughput is so high, how do we think about verification? That's this Other shift that I'm seeing."
「编码不再是瓶颈。……现在,不仅仅是工程师,我们还有设计师、PM,Claude Code 团队中的每个人都在提交代码。……而且产出率如此高,我们如何考虑验证?这是我看到的另一个转变。」

(2026-08 注:这个位移已外溢出工程职能。Anthropic 首位 technical PM Dianne Penn 描述她把月度业务回顾的写作委托给 Claude 后的自我定位:"And I'm more of a reviewer and a verifier of that information."——「而我更多是这些信息的审阅者与核验者」。)

分歧在哪

阵营 A · "coding 已经被解决"——任务被压缩到接近零

Anthropic 的 Boris Cherny 给出了最干脆的说法:

"I think at this point it's safe to say that coding is largely solved. At least for the kind of programming that I do, it's just a solved problem because quad [Claude] can do it."
「我认为现在可以肯定地说,编码在很大程度上已经解决了。至少对于我所做的那种编程来说,这已经是一个解决的问题,因为 quad(即 Claude)能做到。」
Boris Cherny · Head of Claude Code
"I have never enjoyed coding as much as I do today because I don't have to deal with all the minutia."
「我从来没有像今天这样享受编码,因为我不必处理所有的细节。」
Boris Cherny · Head of Claude Code

2026-06 的内部更新:Boris 的直属上级、Claude Code 与 Cowork 团队负责人 Fiona Fung(2026-06-22)给"coding solved"补上了管理层视角。她不否认结论:

"It's lifted the ceiling of what anyone is able to do."
「这打破了任何人能做到的上限。」

但她把"解决之后"重写成一份新的岗位说明书——自主权与问责捆绑出售:

"We say with high agency is also high accountability. So it's all about making sure folks have that freedom to cook. But then it's also like, okay, what's the accountability for it? What's a hypothesis of what you're trying to solve?"
「我们说,高自主权也意味着高责任。所以这都是为了确保人们有自由去创造。但接下来还有一个问题,责任是什么?你想要解决的问题的假设是什么?」

比口号更能说明"solved"边界的是她实际在做的事:她加入 Claude Code 后发现团队缺的恰恰是系统/分布式系统背景的深度专家,于是把招聘改成双轨——"有产品直觉的创造型 builder" + "守住硬核部分的深度系统专家":

"it's all about trust but verify. The models are really good, but there are definitely a lot of areas that still need the verification."
「都是信任但需验证。模型确实很好,但仍然有很多领域需要验证。」

换句话说:在"coding 已被解决"的公司内部,"解决"是按任务类型分层的——生成层交给模型,验证层和深水区仍然明码标价地招人。这等于部分回答了本页"都没说透的"里对 Boris 的追问。配套的流程也换了:六个月路线图三个月就作废,她现在只做 "JIT planning"(月度轻量计划、每周复核)。她还随口给出了一个此前没人提过的代价:

"It could start being a lonely experience because we all started just working with our agents so much. On the Claude Code team recently, we started a pairwise programming lunch."
「这可能会变得孤独,因为我们都开始过多地与我们的代理人合作。最近在 Claude Code 团队,我们开始了一次对编程午餐。」

阵营 B · "瓶颈是人类,不是模型"——工程师*更多*而不是更少

OpenAI Codex 的 Alex Embiricos 反对"coding 被自动化等于工程师变少"的论点:

"Now that we no longer write assembly, like when that change happened, and we moved to higher level languages, did we say coding is automated? Not really, right? We were just able to write much more code. … But every time that's happened, there's been an explosion of demand for the output. And so you need many more people actually to do that kind of work, even if the specific task has changed."
「我们现在不写汇编了——当那个变化发生、我们转向高级语言时,我们会说'编码被自动化了'吗?并不会,对吧?我们只是能写多得多的代码了。……但每次这种事发生,对产出的需求都会爆炸式增长。所以你实际上需要多得多的人来做这类工作,哪怕具体任务变了。」

他那句招牌论断——"人类打字速度和验证工作是 AGI 的关键瓶颈,而不是模型、算力或架构"——出自他此前的推文与写作,访谈中系主持人 Lucas Swisher 当面复述("you said that human typing speed and validation work is the key bottleneck to AGI, not model, compute, or architecture"),并非 Embiricos 在访谈中亲口说出的句子。但他当场认领并展开:

"I think there are multiple bottlenecks, but that's maybe the most sort of clickbaity one. … that's a lot of work to manage these agents and make sure they're always working. … when we look at how often Codex users are using Codex, it's this tens of times range. And I think AI should be helping us tens of thousands of times per day"
「我认为存在多个瓶颈,但这可能是最吸引眼球的一个。……管理这些代理并确保它们始终在工作需要付出很多努力。……当我们观察 Codex 用户使用 Codex 的频率时,大概是几十次的范围。我认为 AI 每天应该帮助我们数万次」

Atlassian 的 Mike Cannon-Brookes 押同一个方向:

"There's no doubt in my mind that we will create far more technology, right? … five years from now we'll have more engineers working for our company than we do today. More software developers working for our companies."
「我毫不怀疑,我们会造出多得多的技术,对吧?……五年后,我们公司雇的工程师会比今天多。为我们公司工作的软件开发者会更多。」
Mike Cannon-Brookes · 20VC: Atlassian CEO

Kevin Weil(前 OpenAI CPO、现负责 OpenAI for Science,2026-06-30)从另一个方向支持"瓶颈是人"——但他给出的解法不是"更多工程师",而是每个人变成多线程操作员。他甚至把"合上笔记本前没给 agent 派活"算作事故:

"I had not gotten a Codex job running. Before I closed my laptop and I was like, shit, I just wasted an hour. … My Codex agent could have been fixing a bug or implementing a feature or doing something for me. … And if you're really good at it, you are not just juggling one job. You've got three or four things running in parallel across different work trees."
「我当时没让 Codex 运行任务。合上电脑后我想,糟了,我刚才浪费了一个小时。……我的 Codex 智能体本可以帮我修复 bug、实现某个功能或者做点别的。……如果你足够擅长,你就不只是在处理一份工作。你可以同时在不同的工作路径上并行运行三四件事。」
"I just think this moment kind of selects for people who are high agency Because you can now create anything that you can think of"
「我认为这一刻会选择那些具有高主动性的人,因为你现在可以创造任何你能想到的东西」

"agency"这个词值得停一下:Fiona Fung(Anthropic)与 Kevin Weil(OpenAI)在相隔八天的两场访谈里各自独立把它当作新的稀缺人才特质——这与 Embiricos 的"人类是瓶颈"是同一枚硬币的两面:模型侧供给近乎无限之后,约束条件全部堆到了人的主动性与验证带宽上。附带一个 B 阵营的实证脚注:主持人 Lenny 当面向 Fiona 指出"按理 AI 应该让工程师更不必要,但你们(和 OpenAI)都在疯狂招工程师"——Fiona 没有反驳,只是把话题引向了下一代怎么培养(见"都没说透的")。

2026-08-24 延长线:阵营 B 拿到了迄今最完整的机制论证——来自放射科的对照组。 Benchmark 合伙人 Eric Vishria(本页第二位 Benchmark 合伙人,与 Ev Randle 同店)用 Geoffrey Hinton 2016 年"该停止培训放射科医生"的著名误判做对照:模型读片能力的判断是对的,结论却错了——聚合训练数据不存在(放射科医生每天读的是全身多模态影像、AI 只攻下单点如胸部 CT)、误读有医疗责任与报销制度绑定、现实于是走进漫长的 co-pilot 期,而效率提升反过来推高了检查总量:

"In the ensuing time, we need more radiologists, not less, because, oh, by the way, everyone's getting more imaging than they used to get because the cost of energy is going down in a Givens paradox kind of way. My point on it is you have someone very, very smart who really understands the capabilities, really understands what's happening, has the right data, but by not thinking of that data in the real world application comes to the wrong conclusion. And that's how I think of the unemployment thing."
「在此期间,我们需要更多的放射科医生,而不是更少,因为,顺便说一下,每个人获得的影像比以前要多,因为能源成本以一种给定悖论的方式下降。我的观点是,你有一个非常聪明的人,真正理解能力,真正理解正在发生的事情,拥有正确的数据,但如果不考虑这些数据在现实世界中的应用,得出错误的结论。我对失业问题的看法就是如此。」

("Givens paradox" 为转写原文,即 Jevons paradox/杰文斯悖论;他紧接着点明软件工程的失业预测"几乎是完全相同的设置"——逐字稿原句 "it's almost the exact same setup"。披露:他投了做放射科 AI 的 New Lantern。)这给 Embiricos 的汇编类比补上了缺失的摩擦项清单:为什么"能力已足够"不会立刻变成"人变少"——数据碎片化、责任归属、co-pilot 过渡期、需求弹性,每一条都是时间的盟友。

阵营 C · "全新编程范式"——不是 coding 变快,是 coding 这个概念被替换

Cursor 的 Michael Truell 不同意 Boris 也不同意 Embiricos——他认为大家都还在用旧词描述新东西:

Cursor 的目标,按 Truell 的说法,是创造一种新型编程、一种构建软件的截然不同的方式:越来越多的工程师会开始觉得自己像"逻辑设计师"——产物不再是数百万行难以理解的代码,而是更真实、更易读、更好导航的东西。(该访谈英文原音、podwise 仅存中文译文,故转述、不作逐字引用。)

Cursor 的 Cloud Agents 团队把同一框架推到团队层:

"We think that over the coming months, the big unlock is not going to be one person with a model getting more done, like the water flowing faster. It will be making the pipe much wider."
「我们认为在接下来的几个月里,最大的突破不是某个人通过一个模型完成更多的工作(像水流得更快一样),而是让管道变得更宽。」
Cursor's Third Era: Cloud Agents · Cursor's Third Era

Karpathy 给这个范式一个最具煽动性的描述——客户身份本身在变:

"Their default workflow of building software is completely different as of basically December."
「他们构建软件的默认工作流程,基本上从十二月起已经完全不同了。」
"The industry just has to reconfigure in so many ways that the customer is not the human anymore. It's agents who are acting on behalf of humans."
「整个行业必须以多种方式重新配置——客户不再是人类,而是代表人类行事的代理。」
"Maybe there's an overproduction of lots of custom bespoke apps that shouldn't exist because agents kind of crumble them up and everything should be a lot more just like exposed API endpoints and agents are the glue of the intelligence that actually tool calls all the parts."
「也许会出现大量本来就不该存在的定制 app 的过度生产,因为 agent 会把它们捏碎——一切都应该更多地暴露成 API endpoint,而 agent 是真正去 tool call 所有这些部分的智能粘合层。」

2026 年年中,阵营 C 得到了两条新证据线——一条来自研究内部,一条来自资本市场。

OpenAI 研究负责人 Mark Chen(2026-06-27)证实 vibe coding 的结构正被原样复制到研究本身——业内已经叫它 "vibe research":

"I think both at OpenAI and at other labs, you're starting to see a lot of the work become mostly orchestration focused, right? Like the researchers coming up with ideas. And the model's great enough to do the implementation execution by itself."
「我认为在 OpenAI 和其他实验室,你开始看到很多工作主要集中在编排上,对吧?就是研究人员提出想法。而且模型足够优秀,能够自己完成实施执行。」

如果"编程这个概念被替换"成立,那么被替换的就不止编程——Mark Chen 描述的正是"研究"这个概念的同构替换(人出 idea 与 taste,模型做执行与编排)。而且他给了本页目前最可证伪的时间表:

"when we look at our kind of three-year roadmap, the end goal that we want to reach is one where The models are just doing end-to-end research and I think a part of that problem is just being able to have the model come up with good taste."
「当我们看我们的三年路线图时,我们想要达到的最终目标是模型能够进行端到端的研究,我认为其中一部分问题就是能够让模型产生良好的品味。」

Benchmark 的 Ev Randle(2026-07-01)给出这场范式替换的商业模型版本——如果 Truell 说的是产物变了、Karpathy 说的是客户变了,Randle 说的是卖的东西本身变了:从卖软件许可证变成卖"随取随用的智能/白领工作":

"It allows a buyer of software to move their mental framing from like, oh, I buy a license to, oh, I buy intelligence on tap or like a white collar activity on tap or some activity that created some economic output for my business on tap via an API or via a software product."
「这使得软件的买家能够将他们的思维框架转变为,哦,我购买的是许可证,变成了哦,我购买的是可随时获取的智能,或者是像是白领活动的可随时获取,或者是通过 API 或软件产品为我的业务创造一些经济产出的活动。」
Ev Randle · Benchmark's AI Bets
"what we've seen in code is going to happen in most white-collar job functions and eventually probably most blue-collar job functions as well."
「我们在代码中看到的情况将会发生在大多数白领工作职能上,并最终可能发生在大多数蓝领工作职能上。」
Ev Randle · Benchmark's AI Bets

范式换了,估值规则也整套倒挂——他称之为"电子表格投资时代的终结",最反直觉的一条:

"if your gross margins are high, that's actually a bad thing. Because AI inference costs a lot of money. And if you have an AI product with high gross margins, that means that no one's using your AI features."
「现在如果你的毛利率很高,那实际上是件坏事。因为 AI 推理成本很高。如果你有一个毛利率高的 AI 产品,这意味着没有人在使用你的 AI 功能。」
Ev Randle · Benchmark's AI Bets

他算的账也值得记下:开发者工具从 SaaS 时代"平均客户约 20 万美元的 line item",正在变成"每个开发者 3.6 万美元/年"、乃至"平均客户 2000 万美元量级的 line item"。另外,他从收入曲线侧标定的拐点与 Karpathy 的"December"说互相印证:

"I don't think we had unbelievable quality coding models until Opus 4.5, until last winter, which is why you saw Anthropx revenue and why you saw Cloud Code go so parabolic, because there was a genuine breakthrough in the usability of those models."
「我认为,直到 Opus 4.5,直到去年冬天,我们才拥有令人难以置信的高质量编码模型,这也是你看到 Anthropic 收入和 Cloud Code 增长那么快的原因。因为在这些模型的可用性上确实有了突破。」
Ev Randle · Benchmark's AI Bets

2026-08 的两条延长线。 其一,Factory 创始人 Matan Grinberg 把"管道变宽"推到尽头,给终局起了个名字——「黑暗工厂」(灯关着、机器自己造机器),并给出本页第二个硬时间表:

"if everyone woke up sick tomorrow, A lot of Claude code usage would be zero because it's all just, hey, Claude code or hey, Codex or hey, Droid, right? I think in 12 to 24 months, like 90% of tokens will be asynchronous tokens."
「如果明天每个人都生病了,很多 Claude 代码的使用就会变为零,因为它全都是,嘿,Claude 代码或者嘿,Codex 或者嘿,Droid,对吧?我认为在 12 到 24 个月内,90% 的令牌将是异步令牌。」
Matan Grinberg · Factory's 'Dark Factory'
"This idea of a dark factory where the lights are off and things are just happening, that is where software development is going."
「这种黑暗工厂的理念,灯光熄灭,事情正在发生,这就是软件开发的发展方向。」
Matan Grinberg · Factory's 'Dark Factory'

值得记下他的自我克制:他承认当下"we're still kind of in like co-pilot mode"(还在副驾模式)——90% 异步是预测,不是现状。其二,「概念被同构替换」的清单还在变长:coding(本节)→ research(Mark Chen 的 vibe research)→ 现在轮到产品管理。Anthropic 首位 technical PM Dianne Penn:

"we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs"
「我们确实会写一些产品文档和 PRD,但团队里有句话叫 “评估即是新版 PRD”」

她同场还有一句押韵的新工作箴言:"Here, you have to sweat the tokens as much as you sweat the pixels."(「在这里,你必须像对待像素一样认真对待令牌。」)——PM 的手感正在从像素移到 token。

2026-08-24 延长线:Karpathy 的"客户不再是人类"拿到了第一个买方侧的具体机制——而且顺手击落了一个流行叙事。 Eric Vishria(Benchmark)先把话说死:

"I'm very dismissive of this idea that people are going to vibe code their own shit and whatever. That isn't the issue at all with SaaS companies. The issue with SaaS companies is the competitive frontier completely shifted."
「我对人们会为自己的事物编写代码之类的想法非常不屑。这根本不是 SaaS 公司的问题。SaaS 公司的问题在于竞争边界发生了完全的变化。」

威胁 SaaS 的不是"人人自建软件"(他不屑的正是 08-23 已退库的 single-use software 叙事的民间版本),而是买方换了物种。他举数据库为例——迁移曾是软件业头号不可为之事,正因如此才有数据库的厚利与粘性:

"Now, you don't have a developer building against the database interface. You have Claude or Codex building against the database interface, one. Two, the beauty of database interfaces is they're very, very well specified. Well, it turns out AI is very good at things that are very, very well specified. Three, agents don't get tired of monotonous work of translating one specification to another. So it turns out that now, all of a sudden, database migration, which used to be the number one thing you would not do in software, is kind of critical."
「现在,您没有开发者在数据库接口上进行开发。您有 Claude 或 Codex 在数据库接口上进行构建。第二,数据库接口的美在于它们非常明确。显然,AI 擅长那些非常明确的事情。第三,代理对单调的工作不会感到疲倦,比如将一个规范转换为另一个。所以现在,数据库迁移突然间变得至关重要,这曾是您在软件中不会做的第一件事,变得相当关键。」

"针对数据库接口写代码的开发者"——这个物种正是数据库公司事实上的客户;如今这个位置上坐着 Claude 和 Codex。这是 Karpathy"客户是代表人类行事的代理"迄今最具体的商业案例:不是产品设计层面的 agents-first(那仍缺席,见"都没说透的"),而是竞争标准层面——粘性护城河失效,赢家标准改写为服务成本与 zero-to-infinity 伸缩。他的结论句式与 Truell 的"新型编程"同构:"It isn't that we don't need databases or that everyone's going to build their own database… What's going to happen is the criteria changed."(「并不是说我们不需要数据库,或者每个人都要建立自己的数据库……会发生的是标准改变了。」)——概念不消失,被替换的是概念的判分规则。

阵营 D · "创作者解放派"——核心变化在 *谁* 能写软件,不在 *怎么* 写

Ryo Lu(Cursor 设计负责人)那期对谈里,"设计师跨过工程师边界"的位移判断其实出自同场嘉宾、a16z 普通合伙人 Jennifer Li 之口(注:此前版本误归为 Ryo 本人):

"I feel like for the first time that design is such an approachable concept and skill set to a lot more people. And it brings together sort of people who have aspirations for design and wanting to build things, wanting to prototype things. Putting beautiful stuff out in the world much, much easier and faster."
「我觉得很多人第一次觉得设计是一个如此平易近人的概念和技能。它将那些渴望设计、想要构建东西、想要制作原型的人们聚集在一起。将美好的东西更快更容易地带给世界。」

Ryo Lu 本人在同场紧接着加的是一个非常重要的限定:

"There needs to be something for the human to specify: What is good? What is right? How I want to do it? If you don't put in that opinion, it will just produce AI slop."
「需要由人来指定:什么是好的?什么是对的?我想怎么做?如果你不放进这层判断,它只会产出 AI slop。」

Kevin Weil 给这个阵营补上了数量级——这是目前语料里对"平权"规模最具体的估算:

"There aren't that many people in the world that know how to program. I don't know, like 30 million people maybe. And you expand it by a couple orders of magnitude, you get an explosion of creativity because lots of people have ideas and sometimes they didn't have any route to actually implement those ideas."
「了解编程的人并不多。我不知道,也许大约 3000 万人。当你将这个人数扩展几个数量级时,你会获得创意的爆发,因为很多人都有想法,有时他们并没有任何途径去实现这些想法。」

他举的例子恰好不是设计师,而是一位小城官员:数据都在、需求清楚、以前雇不起开发者,现在一条 prompt 就能做出当年做不成的市民信息服务。这与上文 Jennifer Li 的设计师叙事互补——平权发生的位置未必在"创意职业"里,更可能在长尾的"从来没被服务过的需求"里。

2026-08 的两个新落点。 OpenAI 把"平权"直接产品化了:核心产品工程负责人 Akshay Nathan 解释为什么把 Codex 并进 ChatGPT Work、而不是留成开发者专属工具——因为职业边界本身在融化:

"that might actually blur the lines between someone who's like only writing code or creating strategy docs or, you know, Planning events or helping with marketing or doing podcasts or whatever, right?"
「这可能会模糊那些仅仅是编写代码或创建策略文档,或者计划活动、协助市场营销、录制播客等等的人的界限,对吧?」

而在 Cursor 自己家里,第一个以"职能"为单位宣布自建工具的不是设计师,而是招聘团队。Head of Talent Adam Ward:

"my one take is like every recruiter should be a talent engineer. Like the things that recruiters now have at their fingertips that they could build in Cursor or other tools. It's really phenomenal. And in earlier companies I was in, we're often beholden to like begging for engineering time to build us tools or functions or features. And now that's like that's in us."
「我的一个观点是:每个招聘官都该是 talent engineer。招聘官现在指尖就有能在 Cursor 或其他工具里构建的东西,这真的很惊人。在我以前待过的公司,我们常常得央求工程排期来给我们做工具、功能或特性。而现在,这能力就在我们自己手里。」

(注意样本偏差:这是 Cursor 自家的招聘团队,天底下最不缺工具氛围的地方;但"不再排队等工程排期"这个机制描述,与 Jennifer Li 的设计师、Kevin Weil 的小城官员完全同构。)

2026-08-24 的两个再落点——外加一个悖论。 其一,"平权"的宣言这次由最大实验室的设计掌门亲口给出。OpenAI 产品设计负责人 Ian Silber(前 Instagram 八年)那期的标题就是立场——"这是历史上做设计师最好的时代":

"I think anybody can be a designer, right? … I think that products that fully understand their users and that make something delightful When anybody can make anything, you know, that's going to become one of the most important kind of like sort of dimensions of a product and that's how people are going to evaluate it."
「我认为任何人都可以成为设计师,对吧?……我认为那些完全理解用户并带来愉悦体验的产品,在任何人都可以创造任何东西的情况下,将成为产品最重要的维度之一,人们也会用这个来评估产品。」

但这期真正的信息藏在开场:主持人 Lenny 的科技职群情绪调查显示,设计师与用户研究员是全维度最不快乐的职群——最焦虑、最疲惫、最不看好前景、最不推荐后辈入行(主持人自述的调查数据,转述)。平权叙事的头号受益人恰是情绪最差的人群,这个悖论 Ian 没有否认,只解释了成因:

"We're unclear right now what is expected of a designer. Someone who might be classically trained in a certain way of working now feeling like, oh my gosh, I need to really change everything."
「目前我们不太清楚设计师的期望是什么。一些可能在某种工作方式上受过经典培训的人,现在感觉哇,我真的需要改变一切。」

他对"角色都在融化"的回答与 Akshay 的"边界模糊"同向、但多一个限定:界限一直模糊、正在更模糊,可大公司里"建造伟大软件需要戴不同的帽子"——职能帽子仍在,只是每顶帽子下的人能做更多别的事(转述其原话 "there are different sort of hats that you need to wear to build great software")。其二,基础设施一线首次为"零代码者上产线"背书。Fireworks CEO 林乔看着平台上长出的海量应用,给出这个时代最干净的门槛描述:

"You have great idea. We want to implement it scaling production. It requires a team of tens of very strong product engineers, PMs working together, multiple quarters to deliver that. And right now, with one person, a few weeks, without understanding how to write a single line of code, you can do that."
「你有很好的想法。我们想要实施它并进行规模化生产。这需要一个由数十名非常强大的产品工程师、项目经理共同合作团队,多个季度才能交付。而现在,靠一个人,几周内,在不懂得写一行代码的情况下,你也能做到这一点。」

注意她讲这话的用意——不是庆祝平权,而是警告平权的代价:门槛塌了,应用层的护城河也跟着塌了("从截图,砰,生成一样甚至更好的 app")。这个反转在暗流一节展开。

Wanaka 创始人张阳 在"解放派"立场内部,画了一条更尖锐的内部分界——平权是有限度的:

最主要其实是看到能让那些不能写代码的人能做出游戏了。
— 张阳 · AI + 游戏 + 社交的新演绎 | 对谈 Wanaka 创始人张阳
我觉得创作能力这件事情它其实是很难真的被平权掉的。
— 张阳 · AI + 游戏 + 社交的新演绎
但是大量的 AIGC 出来的这种普通人的内容,它其实只有一个去向,就是给你的朋友看。……它只对你的朋友价值。
— 张阳 · AI + 游戏 + 社交的新演绎

更凝练的那句"因为他消费的其实是这个关系,消费的不是那个内容本身"出自主持人曲凯之口——是他对张阳前述观点的当场概括,张阳以"对,所以……"接过并展开成上面这段(注:此前版本误把曲凯的概括归为张阳)。

新对撞(2026-08)· Token 经济学:"尽情烧" vs "按回报配给"

这批访谈里几乎每场都绕不开 token 账单——vibe coding 的争论第一次从「能不能」变成「怎么记账」,而两边都有一线数据。

批判 token maxing 的一边。 Factory 的 Matan Grinberg 复盘了这股风潮的成因——董事会问 AI 战略→CTO 把用量写进绩效→人人拿最贵的模型干一切:

"It's this token maxing where People are using like Opus for literally everything like what's the weather NSF Opus? … There are banks that we are working with Where they are spending literally hundreds of thousands of dollars a month on people asking things like, literally, what is the weather?"
「这是代币最大化,在这里人们使用 Opus 几乎所有的事情,比如说天气是什么。……有些银行与我们合作,其中他们每月花费数十万美元让人问诸如天气是什么的问题。」
Matan Grinberg · Factory's 'Dark Factory'

他的解法值得单独看——分层路由的标尺恰好就是 vibe coding 本身:

"the PMs who are like vibe coding dashboards get the same token limits as like the engineers who are building like critical infrastructure. That's probably not the best thing to do. … this part of the org, they're just vibe coding. They can use Gemini Flash."
「那些像 vibe 编码仪表板的产品经理和那些构建关键基础设施的工程师获得相同的代币限制。这可能不是最好的做法。……组织的这个部分,他们只是 vibe 编码。他们可以使用 Gemini Flash。」
Matan Grinberg · Factory's 'Dark Factory'

在他的采购矩阵里,"vibe coding"被归为低风险层配便宜模型、COBOL 配微调模型、关键基础设施配前沿模型多方交叉验证("generate the code with OpenAI, test it with Anthropic, review it with Gemini")——这等于从企业采购侧再次印证了 Fiona 的「按任务类型分裂」:分裂不再只是观察,已经写进了价目表。Elad Gil 把同一套纪律推广成一个指标名:

"I think there's this broader concept of return on invested tokens, like an ROIT kind of metric, which is if you have a certain token budget, who do you give it to and why?"
「我认为有一个更宽的概念——token 投入回报率,类似 ROIT 这样的指标:如果你有一笔 token 预算,你把它给谁、为什么给?」

认为账单焦虑被夸大的一边。 Accel 成长团队拿出了买方调查:

"So we did this survey of developers and we asked them how many of them had a CFO who was telling them to spend less on consumption versus more. And we actually found that seven times more companies are being told to like let it rip and spend more."
「所以我们对开发者进行了一项调查,询问他们中有多少人有 CFO 告诉他们在消费上要减少开支,而不是增加。我们实际上发现,七倍的公司被告知要尽情花费,增加开支。」
Accel 成长团队(说话人转写未具名)· Accel: The Quiet Firm

而为"烧钱"辩护的最认真版本来自 Anthropic 首位 technical PM Dianne Penn(她那期节目标题就带着 token maxing)。主持人 Lenny 转述 Garry Tan 的说法——"每年愿意烧 10 万美元 token 的人,过的是 2028 年人的生活"——她的回应是把账记到另一个科目上:

"It's almost like token spin is more the input. And really, the output is what you described of experimentation. … I will say internally, some of the most creative thinkers, the best like prototypers, do spend a lot of time with Claude, with every new version of a research model that we have."
「这几乎是代币支出更多是输入。而实际输出就是你所描述的实验。……我会说,内部一些最有创造力的思想家,最优秀的原型创造者,确实花了很多时间与 Claude 合作,与我们拥有的每一个新版本的研究模型。」

("token spin" 为转写原文,即 token spend。)

细看之下,这场对撞没有表面那么对立:Accel 说的是总量趋势(全球渗透才刚开始,让它烧),Factory/Elad 说的是单位纪律(谁该用什么模型),Dianne 说的是用途科目(烧在实验上的钱买的是提前住进未来)。真正的信息是:vibe coding 已经大到需要 CFO 和 CIO 出场了。Matan 顺带压了一个可复查的时间点——12 个月内,"每个 CIO 都需要回答:每一颗增量 token 放到哪"(转述)。

2026-08-24 补充:第三种立场入场——账单问题的供给侧技术解。 Fireworks CEO 林乔(Lin Qiao;注意体裁是大会演讲+现场问答,非对谈)从推理基础设施一线给"CFO 出场"提供了第一个具名机制的证据,也把解法从"配给"挪到了"换模型":

"So if we say we need to be careful not scale into bankruptcy, it's not just for startups. It's actually for incumbents. They literally, their CFO is blocking their AI feature launch because of the cost. And post-training is a way to remediate that."
「所以如果我们说需要小心不要在规模扩大中破产,这不仅仅是针对初创公司。实际上也是针对既得利益者。他们的首席财务官因为成本问题而阻止他们推出 AI 功能。而后期训练是解决这个问题的一种方式。」

她的账:后训练出的专用模型可把推理成本压低 5–10 倍、同预算撑 5–10 倍流量(她举 Genspark、Heidi 等客户为例,转述)——"scale into bankruptcy"由此成了这场对撞的第三个词条。把各方摆在一起,光谱现在是完整的:Accel 说总量上尽情烧、Elad/Matan 说单位上按回报配给、林乔说干脆别按前沿模型的价目表烧——自己拥有一个便宜 5–10 倍的专用模型。两个注意:其一,她卖的就是后训练平台,此论亦是其销售论点;其二,"CFO 在拦 AI 功能上线"与 Accel 的"CFO 让尽情烧"(7 倍调查)构成罕见的正面数据对撞——两边都声称看见了 CFO,方向相反。可能的调和是分层:Accel 调查的是开发者工具消费,林乔说的是 incumbent 把 AI 功能铺给全量客户的推理账单——前者是研发科目,后者是 COGS。

暗流 · "AI slop / 过度生产"是反复出现的反音

Karpathy 担心 bespoke apps 的过度生产,Ryo Lu 担心没有人类判断会出 AI slop,张阳 担心普通人创作出来的 AIGC 在陌生人之间无消费价值——三人没有协调过,但都在用近义词描述同一种风险:工具的供给增长远超人类筛选能力。这条线在主流"AI 让人变超人"的叙事下贯穿,却没人正面接住。

"It's almost like I do this because the AI really knows these things well because it has seen it a lot around the internet when they're used for training. So it is really good at composing patterns that exist."
「我这样做是因为 AI 非常了解这些东西——它在训练数据里见过很多。所以它非常擅长组合那些已经存在的模式。」

Ryo 这句话夹在他对 AI 设计的乐观叙述里,但它本质上是 "AI 在做组合,不在做创造" 的一种自承——这跟阵营 A 的"coding is solved" 不在同一种"解决"上。

本次更新给这条暗流添了两个新落点。其一,Ryo 的"taste 论"拿到了跨实验室背书——OpenAI 研究负责人 Mark Chen 在描述 vibe research 时主动承认:

"And that's why you still need the researchers coming up with the ideas. … It's going to be hard to teach the models good taste. We noticed that."
「这就是为什么你仍然需要研究人员提出想法。……教会模型良好的品味将会很困难。我们注意到了这一点。」

Ryo(设计判断)、Fiona(验证)、Mark Chen(研究品味)——三个不同职能、两家对头实验室,指向同一个残余人类职能:判断什么是好的。这让 AI slop 的定义更清晰了:slop 不是模型能力不足的产物,而是人类判断缺席的产物。

2026-08 批次把这条证据线从三条加到五条,且新增两条全部来自两家实验室的产品侧。 OpenAI 的 Akshay Nathan 与 Anthropic 的 Dianne Penn 各自补上产品职能的版本:

"Bottlenecks becomes sort of like ideas and taste, I guess. I think because anyone can build now, I think it really is the era of bottoms-up ambition."
「瓶颈变得有点像想法和品味,我想。我认为因为任何人现在都可以构建,所以这确实是自下而上的雄心时代。」
"I think judgment is one is an area where it's accumulation of so much nuance and so much experience. And these systems haven't experienced as much as humans have. And so I think that hard earned, like judgment is a a area for for product leaders and just generally will continue to be really critical."
「我认为判断是一个积累了如此多细微差别和经验的领域。而这些系统的经历远不及人类。所以我认为经过艰苦努力的判断对于产品领导者来说是一个关键领域,通常也将继续特别重要。」

(顺带一个词汇学观察:Akshay 那场的主持人问的是"你们缺的是设计师,还是 slop cannons(slop 大炮)?"——slop 已经从批评词变成了行业自嘲的日常词汇。)

2026-08-24 批次再加三条,证据线增至八条——且第一次出现内部分叉。 其一,Ian Silber 给"判断是残余人类职能"补上了此前一直缺失的机制解释,它与 Ryo 的"AI 擅长组合已存在的模式"恰是同一枚硬币的正反面:真正的新东西没有训练数据。

"take the iPhone, we're designing for multi touch. And that opened up like this whole new interaction paradigm. And there's no training data for that, right? … humans are able to kind of do a really good job kind of evaluating, you know, what good is … there's a lot of reasons why Having that human in the loop is gonna be so important"
「以 iPhone 为例,我们在为多点触控设计。这开启了一个全新的交互范式。而且对此没有训练数据,对吧?……人类能够很好地评估什么是好的……有很多原因说明,保持人类在循环中是非常重要的」

(他此前一句刚承认"AI 已经是一个不可思议的产品设计师"——所以这不是防御性发言,而是给"solved"划界:语音、agent 交互这些正在被设计的范式,恰好都是没有先例的。)其二,投资侧第一次把三种职能直接溶解成三种能力,taste 位列其中。Eric Vishria:

"I think what there really are, are people who understand customer problems. People who have taste and people who understand the jagged edge of AI capabilities and are curious about. Those are the three things. … it doesn't really matter if you are an engineer or a product manager or designer, but if you have taste and understanding of customer problems and understanding the jagged edge, you're going to do well."
「我真正认为的是什么,是理解客户问题的人。有品位的人,以及理解 AI 能力棱角的人,并对此感到好奇。这就是三件事。……如果你是工程师、产品经理还是设计师并没有真正关系,但如果你有品位,对客户问题的理解,以及对棱角的理解,你会做得很好。」

其三是那道分叉——对这条证据线既是加固也是挑战。Fireworks CEO 林乔的整场演讲只有一个论点:taste 不必留在人脑里,可以被后训练进模型权重,变成偷不走的资产:

"post-training allows you to basically encode, codify your unique taste Into a model that no one can steal from. Because it's very easy to clone and copy application as is, as you all know, right? From coding agent, it's very easy. So from screenshot, boom, generate the same or even better app. So that's very scary that application itself. Is kind of the mode is being reduced."
「后训练让你基本上能够将你独特的品味编码成一个无人可以窃取的模型。因为像大家所知道的那样,克隆和复制应用程序是非常简单的,对吧?从编码代理来说,这很容易。所以从截图开始,生成相同或甚至更好的应用程序。所以这让人非常害怕,应用程序本身的情况。这种模式正在被削弱。」

("the mode is being reduced" 为转写原文,指 moat/护城河在变薄。)这与 Mark Chen 的"教会模型良好的品味将会很困难"表面冲突、实则切开了 taste 这个词:产品性偏好(什么输出算好——可以用偏好数据与奖励函数编码;她点名 Cursor 的 Composer 与 Factory 的安全模型为实例,恰好都是本页已有的声音)与开放式品味(提出好问题、发明没有先例的东西——Ian 的"没有训练数据"、Mark Chen 的"学不会"说的是这个)。前者正在变成可交易资产,后者仍是人类残余职能。她自己的话也自带这条边界:她批评创始人们的评估大多是 vibe evaling——看一眼结果"感觉对不对"——而这恰恰应该被系统化掉:"Actually, this is judgment. You put your judgment of any result and decide whether it's good or not. And that judgment you convert into systemic evolution."(「实际上,这就是判断。你对任何结果做出判断,并决定它是好还是不好。而这个判断转化为系统性的演进。」;"systemic evolution" 为转写原文,即 systematic evaluation)——判断不消失,而是被要求从氛围固化为代码,这与 Dianne 的"evals are the new PRDs"在两家公司异口同声。(立场注记:林乔卖的就是后训练平台,"taste 可编码"同时是她的销售论点。词汇学续报:vibe coding → vibe research → 如今有了 vibe evaling——vibe 前缀继续繁殖,且第一次以贬义登场。)

其二,Kevin Weil 给这条暗流提供了第一个市场化出口的猜想——如果 slop 泛滥,"人造"会变成溢价标签:

"We'll end up with like, because now you still, people still value custom made furniture. And then we'll end up with like bespoke websites. Like this was done by a human."
「我们最终会变成这样,因为现在人们仍然重视定制家具。然后我们最终会有定制网站。就像这是由人类完成的。」

这与张阳的"普通人的 AIGC 只有一个去向——给你的朋友看"(主持人曲凯当场概括为"消费的是关系,不是内容本身",张阳认同)意外接上了:两人都在说,当内容供给无限时,稀缺性会迁移到内容之外的东西上(手作身份 / 熟人关系)。但注意这仍是猜想——没有任何一方给出消费侧数据。

其三(2026-08 新增),这条暗流迎来了一个珍贵的反例——而它反过来支撑了上面的定义。在恶意软件圈,vibe coding 不产 slop、反而提质。Socket 创始人 Feross Aboukhadijeh 谈第一批 NPM 自传播蠕虫(被问"作者大概率用了 AI 吧?"):

"Almost certainly, yes. And there's been, you know, that malware, I think we have pretty good reason to believe that it was vibe coded. There's been one of the threat groups actually posted their You know, open source to their kind of vibe-coded toolkit for others to use to be able to do this."
「几乎可以肯定,是的。而且那个恶意软件,我们有相当充分的理由相信它是 vibe coded 的。有一个威胁组织甚至把他们那套 vibe-coded 工具包开源发布,供其他人拿去干同样的事。」
Feross Aboukhadijeh (Socket) · The Reality of AI-Powered Cyberattacks

同场主持人(转写标注 Speaker 3、未具名)当场把这个反转说破:

"And malware authors were never really great coders. You probably don't realize this, right? So, if the code starts looking better, it's probably vibe-coded, right? It's sort of the opposite of what you think of vibe-coding typically."
「而恶意软件作者从来都算不上好的程序员。你可能没意识到这一点,对吧?所以,如果代码开始变好看了,那多半是 vibe-coded 的。这正好和你通常对 vibe-coding 的印象相反。」
同场主持人(未具名)· The Reality of AI-Powered Cyberattacks

这两句合起来是对"slop = 判断缺席的产物"最干净的一次检验:同一种工具,放在动机极强、目标函数极清晰的作者手里——哪怕是罪犯——产出质量不降反升。slop 从来不是 AI 的属性,是使用者意图与判断的属性。副作用是它也在张阳的"创作能力难以被平权"上撕开了一道口子:至少在"写出更好的恶意软件"这个黑暗角落,平权是彻底的——Feross 同场还提到,攻击载荷现在常常"actually prompts"(就是提示词本身)、借开发者本机装着的 AI CLI 工具借力穿过传统 EDR 检测(转述)。

都没说透的

"all of a sudden, these Agents are making decisions about downstream workflows and downstream tool creation. And we just sort of said, we have to invest in every single company that is going to be in this flow and every single company, more importantly, that is like the choke point that's metering out these decisions."
「突然间,这些代理开始对下游工作流和下游工具创建做出决策。我们就这样说,我们必须投资每一家将进入这个流程的公司,更重要的是,投资那些是在这些决策中起到瓶颈作用的公司。」
Accel 成长团队(说话人转写未具名)· Accel: The Quiet Firm

但注意:这仍是从采用曲线倒推出的间接证据,agents-first 的产品设计实操依旧缺席——问题本身继续开放。)(2026-08-24 更新:Vishria 的数据库案例(见阵营 C)把证据再推近一步——不再是从采用曲线倒推,而是直接指认"针对数据库接口写代码"这个客户位置上已经坐着 Claude/Codex,数据库赢家的判分标准已被改写。但 agents-first 的产品设计实操依旧无人展示——问题保持开放。)

"But I wonder if it's for software engineering, it's almost like you go more towards a fellowship or apprenticeship program."
「但我想知道对于软件工程,是否更倾向于走学徒或见习项目的方向。」

我的看法

判断(不是事实):这四个阵营其实在讨论*不同时间尺度上的同一件事*。1–2 年看:阵营 B(更多工程师、生产力放大)是对的,因为人类验证瓶颈短期内不会消失;3–5 年看:阵营 C(编程概念被替换)逐渐成立,因为"管道变宽"会先于"消除"出现;长期:阵营 A("coding solved")和阵营 D(人人创作)会一起发生,但AI slop 这个反音会决定它们到达不到大众市场——这是最被低估的变量

我对这个判断的把握:中等。对前两段(短期工程师变多、中期范式位移)信心较高;对长期判断信心明显较低,主要因为"消费侧筛选能力"这个变量在所有访谈里都被绕开了,没法做出有数据支撑的判断。

2026-07-16 追记:四篇新访谈没有动摇这个框架,反而把它拧得更紧了。(1)阵营 A 和 B 正在合流成同一句话——"生成被解决了,验证没有"——这从 Anthropic 内部(Fiona)和 OpenAI 产品侧(Embiricos,更早)两头得到确认;(2)"taste + 验证"作为人类残余职能,现在有三条互相独立的证据线(Ryo/设计、Fiona/工程、Mark Chen/研究),这是本页目前最坚固的一个判断;(3)中期阵营 C 的证据在加厚:范式替换已经外溢到研究本身(Mark Chen 的 vibe research)和商业模型(Randle 的 intelligence on tap)。因此对前两段判断的信心由"较高"上调为"高";长期判断的信心维持较低不变——消费侧筛选能力依旧没有任何人给出数据,Kevin Weil 的 bespoke 溢价只是又一个未经验证的均衡态猜想。

2026-08-10 追记(2026-08-23 退库清理后按现存语料修订计数):本批新访谈没有动摇四阵营框架,但给出了一个新的阶段信号:讨论重心正从"范式是否成立"移向"范式怎么记账"(token 经济学对撞、CIO 配给、ROIT)。我把这读作短期判断(阵营 B)的强化证据——当一件事开始需要 CFO/CIO 为它设预算科目,它就不再是实验,而是常设产能。中期(阵营 C)多了一个比 Mark Chen 三年路线图更近的可证伪锚点:Matan 的"12–24 个月内 90% token 异步化",若兑现,"管道变宽"就从隐喻变成计量事实。"taste/判断是残余人类职能"的证据线增至五条且新增两条全部来自两家实验室的产品侧——这仍是本页最坚固的判断,而且 vibe-coded malware 这个反例反而让它更精确了:判断不是"人类还剩下的能力",而是"AI 输出质量的决定变量"——作者在乎什么,产出就像什么。长期判断(slop 决定大众市场天花板)信心仍低,但方向上多了一个间接支撑:连犯罪团伙都证明了目标函数清晰的作者用 vibe coding 会提质,那么大众市场 slop 泛滥的根源就更明确地指向"大多数创作者没有清晰的目标函数"——这与张阳的"熟人关系是唯一去向"互洽。短期 B、中期 C 的框架不变。

2026-08-24 追记:三篇新访谈(OpenAI 设计掌门、Benchmark 二号合伙人、Fireworks CEO)没有动摇框架,但在三个位置把判断拧紧或切细了。(1)本页最坚固的判断——"taste/判断是残余人类职能"——证据线从五条加到八条,但 taste 这个词第一次被切开了:产品性偏好可以被后训练编码进权重、变成资产(林乔,且 Cursor/Factory 已在做),开放式品味仍无人声称可教(Ian 的"没有训练数据"机制 + Mark Chen 的"学不会")。我把这读作判断的升级而非削弱:残余人类职能正在收窄为"在没有训练数据的地方决定什么是好的"——范围更小,但更硬、更可辩护。(2)短期判断(阵营 B)拿到迄今最好的机制论证:放射科对照组把"能力≠替代"拆成数据碎片化、责任、co-pilot 期、需求弹性四个可检验的摩擦项——这比"历史上每次都变多"的归纳论证质量高一个等级。(3)一个此前没人报过的暗数据:平权叙事的头号受益人(设计师)是全维度情绪最差的职群(Lenny 调查)。如果红利与焦虑同源——角色说明书被撕掉——那么阵营 D 的"解放"与阵营 B 的"人是瓶颈"可能是同一件事的白天与黑夜。另记一笔同源风险:Vishria 与 Ev Randle 同为 Benchmark 合伙人,且 Fireworks 正是 Benchmark 的持仓(主持人当面点名"Sierra and Fireworks")——本批"范式已成立"的声音里资本利益浓度偏高,下批取材应刻意找空头或用户侧败例。短期 B、中期 C 不变。

还想知道什么

取材