目录
「Scaling 撞墙了」「there is no wall」「预训练已死」「再堆一个数量级就行」——2024 年 11 月到 2026 年 7 月,这四种句子在同一个行业里同时流通,说它们的人包括这个领域最贵的一群大脑。本篇不问标度律本身对不对(那是神经标度律篇的题),而审一场有确切日期的活论战:2024Q4 起爆的「预训练 Scaling 撞墙/已死」讣告,与「只要继续堆算力、AGI 在望」的命定外推,双向同审。
去重与分界声明:本库已有神经标度律的第一性原理篇(幂律拟合的认识论与四派理论会聚,本篇只取结论不重做)、技术奇点篇(命定外推末端的审计;Ilya「pre-training…will unquestionably end」在该篇作反讽锚出现,本篇把它放回论战时间线原位并做拼接引文考据,两篇零重叠)、LLM 推理篇(RLVR/test-time 机制细节,本篇只用其结论性数字)、AI 就业篇(Acemoglu TFP 账,本篇只定位不重复)、涌现幻象篇(Schaeffer mirage 论证)。本篇管的是「墙叙事」本身的证据链:谁说了什么、数字是否兑现、钱往哪走。
结构胎记=单轴放缓观察冒充全线终审(讣告跳)× 经验规律冒充物理定律(名号跳)× 新轴兑现冒充墙从未存在(反转跳)。三跳各有原话承重:Bloomberg 把「三家公司收益递减」写成行业终审;Amodei 与 Nadella 两位最大押注者都自承 scaling laws 是 empirical regularities 而非物理定律;而 2025–2026 年「撞墙论破产」的反转叙事,又把推理时计算的兑现读成讣告派全错——讣告派的核心观察(预训练单轴增益放缓)恰恰是真的,并被 OpenAI 自己的系统卡自白确认。
母裁决七层硬度光谱:① 唯象地基真而分层(loss 层幂律平滑可预测是 Kaplan 一手原话,但 Kaplan/Chinchilla 均未经同行评审、下游指标可预测性系统性退化、「20 tokens/参数」口诀不在 Chinchilla 原文);② 叙事考古最硬(时间线与当事人原话可逐字复核;「撞墙」一词是媒体框架不是当事人措辞,网传 Ilya 拼接引文是传输层病);③ 数据墙真而被缓冲(公共文本存量有限、Epoch 中位估计 2028 耗尽,但多 epoch/合成数据/多模态三层缓冲,Epoch 自己 2025 年判「unlikely to run out of data」);④ 代际证据双向各得一分(Orion 不及预期真、GPT-4.5 系统卡自白「not a frontier model」真;但 2025–2026 各实验室代际持续推进、新基准被快速攻陷也真);⑤ 推理时新轴真兑现而有天花板迹象(o1/o3 官方数字、R1-Zero 15.6%→77.9% 真;ScaleRL 发现 RL 收益是 sigmoid 有渐近线、2026 年 overthinking 边际递减论文、ARC-AGI-2 起步个位数也真);⑥ 经济账双向同真(撞墙论后 18 个月四大厂 capex 反手自约 $500B 加到约 $700B/年,但 OpenAI 收入缺口未闭、自我把算力承诺从 $1.4T 砍到约 $600B);⑦ 双向裁决:讣告未立、命定未立——讣告派赢在观察(单轴放缓真)输在终审(全线停滞未发生),命定派赢在续命(新轴真兑现)输在外推(平滑可预测失守、经济账未闭、天花板迹象已现)。
〇 母裁决·七层硬度光谱
| 层 | 在说什么 | 认识论地位 | 软在哪类诉求 |
|---|---|---|---|
| ① 唯象地基·硬而分层 | Kaplan 2020 幂律三因子+「平滑可预测」;Chinchilla 2022 等比例律 | loss 层真(七数量级幂律、跨架构弱相关),但两篇均未过同行评审;Chinchilla 被 Epoch 复现抓出统计报告瑕疵却反而背书其结论 | 外推诉求:Bahri 证明指数依赖数据流形维度非普适常数;BNSL「typically (but not always!)」;Schaeffer 2024 证明下游指标可预测性系统性退化——「平滑可预测」到能力层就失守 |
| ② 叙事考古·最硬 | 2024-11 引爆周到 2026-07 的完整时间线与逐字原话 | 逐字可复核(The Verge 现场记录、BNN 全文转载、官方财报电话会转录、官方播客逐字稿) | 终审诉求:「撞墙」一词出自 Bloomberg/Platformer/Marcus 的框架,The Information 引爆原文无 wall 一词;网传 Ilya「三合一」引文是多句拼接 |
| ③ 数据墙·真而被缓冲 | 公共人类文本存量有限;Epoch 估计 2028 中位耗尽 | 存量估计有 5 倍不确定带(510T 原始→100T 质量调整→320T 重复调整);v1(2024 前耗尽)已被 v2(2026–2032)自我修正过一次 | 终审诉求:多 epoch(4 epoch 近似新数据)、合成数据(phi-4 占 40% 配比)、多模态三层缓冲;Epoch 2025《AI in 2030》自判「unlikely to run out of data」 |
| ④ 代际证据·双向各得一分 | Orion 不及预期 vs 2025–2026 持续推进 | 两边都是硬数字:OpenAI 系统卡自白「GPT-4.5 is not a frontier model」;GPT-5→5.1→5.2→5.4、Claude 3.7→4.6、Gemini 2.5→3 代际节奏未停 | 两种单向诉求:拿 Orion 宣判全线停滞(随后 18 个月被证伪),或拿新代际宣布讣告派全错(单轴放缓观察本身是真的) |
| ⑤ 推理时新轴·真兑现/有天花板迹象 | o1→o3→R1→各厂 thinking 模型;test-time compute 成为第二增长轴 | 官方数字可核(o1 AIME 74.4% 附录 A;o3 ARC-AGI 75.7%/87.5%;R1-Zero 15.6%→77.9%) | 永动诉求:ScaleRL 实测 RL 收益是 sigmoid 有渐近天花板;2025-12 大规模比较「no single TTS strategy universally dominates」;2026-04「marginal returns diminish substantially…abandoning previously correct answers」;ARC-AGI-2 起步个位数 |
| ⑥ 经济账·双向同真 | capex 约 $700B/年 vs 收入缺口未闭 | 一手财报数字(MS CY2026 ~$190B、Google 2026 指引 $175–185B、Meta 上调至 $125–145B、Amazon ~$200B) | 两种单向诉求:「市场用钱投票没撞墙」(但 OpenAI 自砍承诺 $1.4T→$600B、Stargate JV 停摆);「泡沫必破」(Covello 2024-06 预言 12–18 个月热情消退,到期未兑现) |
| ⑦ 双向裁决·讣告未立/命定未立 | 讣告派四连错 × 命定派四连软 | 中间派定调最接近落点:Lambert「trade-offs among different types of scaling are real」 | 词义诉求:双方用 2023 年的「scaling」定义打 2026 年的仗——这个词已从单轴裂成预训练/后训练/推理时/智能体四轴 |
一、唯象地基层:「平滑可预测」的出处与边界
「撞墙」论战双方都引用同一组文献当预测基准:Kaplan 2020 说性能随规模「平滑可预测」,讣告派说这条曲线弯了,命定派说它没弯。所以第一层先审这组文献本身到底承诺了什么。
1. Kaplan et al. 2020:幂律三因子与那句被双方各取所需的「无可偏离」
Kaplan et al. 2020《Scaling Laws for Neural Language Models》arXiv:2001.08361(arXiv 仅 v1,未经会议正式发表)摘要逐字:「The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude. Other architectural details such as network width or depth have minimal effects within a wide range.」以及「Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.」(主笔 arXiv 摘要页亲核。)
论战的预测基准在 §1「Summary of key findings」(主笔 PDF 全文亲核):
- 「Smooth power laws: Performance has a power-law relationship with each of the three scale factors N, D, C when not bottlenecked by the other two, with trends spanning more than six orders of magnitude (see Figure 1). We observe no signs of deviation from these trends on the upper end, though performance must flatten out eventually before reaching zero loss.」
- 「Taken together, these results show that language modeling performance improves smoothly and predictably as we appropriately scale up model size, data, and compute.」
指数口径(§1, Eq. 1.1–1.3):L(N) 指数 αN≈0.076、L(D) αD≈0.095、L(Cmin) αCmin≈0.050。注意论文自己区分两个算力口径——§6 汇总表同时给 αC=0.057(原始算力 C)与 αCmin=0.050(最优分配算力),正文明写「it is the trend with Cmin that should be used to make predictions」。后世大量「每 10 倍算力换多少 loss」的二手引用把两个口径混用。
三个必须带进报告的限定:(a) 这些律描述的是 pretraining loss(交叉熵),不是下游能力;(b) Kaplan 自己留了退路「performance must flatten out eventually before reaching zero loss」——「无可偏离」严格限定在观测窗内;(c) 论文未经同行评审,这一点同样适用于 Chinchilla(Rosenfeld/BNSL/Sorscher/Muennighoff/Bahri/Hägele 均有正式发表载体,唯独这两篇最常被当「定律」引用的没有)。
2. Chinchilla:等比例律、「4× more more data」[sic] 与一个口诀的考据
Hoffmann et al. 2022《Training Compute-Optimal Large Language Models》arXiv:2203.15556(arXiv 仅 v1,未经会议正式发表)摘要逐字(主笔亲核):「We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.」「…for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.」以及「Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks.」摘要原文即写作「4× more more data」〔sic,”more” 重复,arXiv 页面与 v1 PDF 均如此〕。
与 Kaplan 的分歧(§1 逐字):「given a 10× increase computational budget, they suggests that the size of the model should increase 5.5× while the number of training tokens should only increase 1.8×. Instead, we find that model size and the number of training tokens should be scaled in equal proportions.」
「20 tokens/参数」口诀考据(本库首次立案):该短语在 Chinchilla 原文中不存在逐字形式——它是从 headline 配置 70B/1.4T(恰为 20:1)与 Table 3 各行(≈20–22,如 400M→8.0B、175B→3.7T、280B→5.9T)派生的二次文献总结。且论文内部口径不一:Table 3 给 280B→5.9T,§3.4 正文给「A 280 billion Gopher-like model…should be trained on 6.8 trillion tokens」。引用口诀时应写「Chinchilla 的 headline 配置对应 20:1」而非「Chinchilla 提出 20 tokens/参数定律」。
Epoch 复现的双向价值:Besiroglu, Erdil, Barnett & You 2024《Chinchilla Scaling: A replication attempt》arXiv:2404.10102 摘要逐字(主笔亲核):「We find that the reported estimates are inconsistent with their first two estimation methods, fail at fitting the extracted data, and report implausibly narrow confidence intervals–intervals this narrow would require over 600,000 experiments, while they likely only ran fewer than 500.」——这是讣告派可用的「地基文献统计有瑕」证据。但同一篇复现的落点是:「In contrast, our rederivation of the scaling law using the third approach yields results that are compatible with the findings from the first two estimation procedures described by Hoffmann et al.」,且其 §A.2 自报「our point estimates imply that 25.6 tokens per parameters is optimal」。批评性复现、整体性背书:Chinchilla 的统计报告有瑕,等比例律本身反而站住了。
3. 批评谱系:幂律不是普适形式,「可预测」到下游就退化
按论证类型分四摞(全部一手逐字,主笔抽查亲核其中摘要):
唯象限定摞。Caballero et al. 2022《Broken Neural Scaling Laws》arXiv:2210.14891(ICLR 2023;注意该文有 17 个 arXiv 版本)§1 逐字:「The upstream/in-distribution test loss typically (but not always!) falls off as a power law with increasing data, model size and compute. However, the downstream/out-of-distribution performance…are often harder to predict, sometimes exhibiting inflection points…and non-monotonic behaviors.」Xiao et al. 2024 arXiv:2409.15156 提出「scaling law crossover」——两条标度曲线可在某规模处相交,小尺度有效的方法外推到大尺度失效;§5.4 逐字:「naively extrapolating scaling laws for model comparison can be un-reliable and lacks a solid theoretical foundation.」
理论解释摞。Bahri et al.《Explaining Neural Scaling Laws》arXiv:2102.06701(PNAS 2024 正式版)摘要逐字:「We identify variance-limited and resolution-limited scaling behavior for both dataset and model size, for a total of four scaling regimes.」其机制(§1.3)是模型把数据流形切成更小区域,指数 α∝1/d——指数由数据流形内蕴维度决定,不是普适常数。这从理论上封杀了「一条幂律走天下」的读法。
逃离幂律摞。Sorscher et al. 2022《Beyond neural scaling laws》arXiv:2206.14486(NeurIPS 2022)§1 判词逐字:「Such power law scaling has motivated significant societal investments in data collection, compute, and associated energy consumption. However, power law scaling is extremely weak and unsustainable.」——并实证数据剪枝可把误差对剪枝后数据集大小的关系打成指数级改善(限定:依赖高质量剪枝度量,ImageNet 规模已吃力)。Kumar et al. 2024《Scaling Laws for Precision》arXiv:2411.04330 给出更怪的反例:训练数据越多,后训练量化退化越重,「eventually making additional pretraining data actively harmful」(摘要逐字)。
方法学与效度摞。Hägele et al. 2024 arXiv:2405.18392(NeurIPS 2024;⚠️ 本报告任务书曾误植为 2405.17818,那是一篇高光谱图像融合论文)证明 Chinchilla 依赖的 cosine schedule 会「underestimates the model performance during training」,constant-LR+cooldown 可用可复用 run 拟合标度律——修正的是标度律实验方法学。Alabdulmohsin et al. 2022 arXiv:2209.06640(NeurIPS 2022)更早指出内插拟合好≠外推准,应以 extrapolation loss 为准。Schaeffer et al. 2024 arXiv:2406.04391(Stanford 等)直接打在论战七寸上:对 MMLU 等选择题基准,「Downstream performance is computed from negative log likelihoods via a sequence of transformations that progressively deteriorate the statistical relationship between performance and scale」(作者主页自述)——loss 层的平滑可预测,经过指标变换链到下游就系统性退化。Diaz & Madaio《Scaling Laws Do Not Scale》arXiv:2307.03201(FAccT 2024)则是规范性批评:「models may not, in fact, continue to improve as the datasets get larger — at least not for all people or communities impacted by those models」(摘要逐字)——论证类型是测量效度,不是唯象证伪,引用时勿混。
优先权注:幂律现象的早期系统化是 Rosenfeld et al. 2019 arXiv:1909.12673(ICLR 2020,早于 Kaplan 四个月),其 §1 自陈「the first work that provides simultaneously: A joint functional form of the generalization error landscape—as dependent on both data and model size」;更早的幂律观察可溯至 Hestness et al. 2017。
4. 第一层裁决
loss 层的幂律与「平滑可预测」是真锚——它是双方共同的预测基准,也被此后五年各家 frontier 训练反复利用。但三层软化在论战爆发前就已就位:(a) 两篇地基文献未经评审,统计报告被 Epoch 复现抓出瑕疵;(b) 幂律不是普适函数形式(BNSL/crossover/Bahri 机制);(c) 最关键的是 Schaeffer 2024 证明的:可预测性沿「loss→指标变换→下游能力」链条系统性退化。双方后来在 2024Q4 争论的「曲线弯没弯」,量的根本不是同一匹布——讣告派量的是下游代际增益,命定派量的是 loss 与基准分,而第一层文献只承诺过 loss。
二、叙事考古层:「撞墙」这个词是谁说的
1. 引爆周(2024-11-09 → 11-14):三篇报道与一条四个词的推文
2024-11-09,The Information(Stephanie Palazzolo、Erin Woo、Amir Efrati,《OpenAI Shifts Strategy as Rate of ‘GPT’ AI Improvements Slows》,原文页(付费墙))。原文导语逐字:「The rate of improvement for the basic building blocks underpinning them appears to be slowing down, though.」经 Platformer 2024-11-14 块引用核到关键段:「While Orion’s performance ended up exceeding that of prior models, the increase in quality was far smaller compared with the jump between GPT-3 and GPT-4…」「Orion performs better at language tasks but may not outperform previous models at tasks such as coding.」注意:引爆原文通篇无 “wall” 一词——「撞墙」框架是随后 48 小时内由 Bloomberg 定调、Platformer 标题(”AI companies hit a scaling wall”)与 Gary Marcus 加上的。
2024-11-11,Reuters(Krystal Hu、Anna Tong,《OpenAI and others seek new path to smarter AI as current methods hit limitations》,原文页)。两条承重逐字(经 Platformer 块引用核实):Ilya Sutskever 称预训练 scaling 的结果「have plateaued」;其直接引语:「The 2010s were the age of scaling, now we’re back in the age of wonder and discovery once again. Everyone is looking for the next thing. Scaling the right thing matters more now than ever.」同稿引 OpenAI 研究员 Noam Brown(TED AI 2024-10):「It turned out that having a bot think for just 20 seconds in a hand of poker got the same boosting performance as scaling up the model by 100,000x and training it for 100,000 times longer.」(另经 VentureBeat 2024-10-23 独立核到同句。)
2024-11-13,Bloomberg(Rachel Metz 等四人,《OpenAI, Google and Anthropic Are Struggling to Build More Advanced AI》,原文页(付费墙);主笔经 BNN Bloomberg 全文转载亲核)。定调句逐字:「After years of pushing out increasingly sophisticated AI products at a breakneck pace, three of the leading AI companies are now seeing diminishing returns from their costly efforts to build newer models.」细节:Orion「did not hit the company’s desired performance」、Gemini 新代际「not living up to internal expectations」、Claude 3.5 Opus 时间表滑移且「performed better on evaluations than the older version but not by as much as it should, given the size of the model and how costly it was to build and run」。引语两则:Hugging Face 首席伦理科学家 Margaret Mitchell「The AGI bubble is bursting a little bit,」;Bentley 大学数学副教授 Noah Giansiracusa「We got very excited for a brief period of very fast progress…That just wasn’t sustainable.」同文转引 Amodei 前一天播客原话(见下条)。
2024-11-14,Sam Altman 推文:「there is no wall」——全推仅此一句(小写)。X 反爬未能直取原推,Platformer 2024-11-14 与 VentureBeat 2024-12-01 两家均按全推形式引用,未见更长版本。Platformer 同文评注:「It was the sort of gnomic pronouncement that Altman has increasingly favored this year…OpenAI had declined to comment on the reports above, but here was Altman appearing to deny them — or if not deny their individual claims, at least to deny their implications.」
2. 当事人接招:四位掌门人的同周表态
Dario Amodei(Lex Fridman #452,2024-11-11,官方逐字稿)。被 Bloomberg 转引、主笔经 Bloomberg 全文亲核的一段:「People call them scaling laws. That’s a misnomer. They’re not laws of the universe. They’re empirical regularities. I am going to bet in favor of them continuing, but I’m not certain of that.」逐字稿中的补充(时间戳 00:06:34):「But I’ve seen the movie enough times, I’ve seen the story happen for enough times to really believe that probably the scaling is going to continue, and that there’s some magic to it that we haven’t really explained on a theoretical basis yet.」(00:15:51)「…I would bet against it, but it’s definitely possible, is we simply run out of data.」注意:行业最大的 scaling 押注者亲口否认「定律」名号——这是本篇名号跳的一手自证。
Jensen Huang(英伟达 FY25Q3 财报电话会,2024-11-20,Motley Fool 官方转录):「Our foundation model pretraining scaling is intact, and it’s continuing. As you know, this is an empirical law, not a fundamental physical law. But the evidence is that it continues to scale.」同场首次给出多轴框架:「it’s not enough…we’ve now discovered two other ways to scale…And so, we now have three ways of scaling, and we’re seeing all three ways of scaling.」
Satya Nadella(BG2 播客,2024-12-16 发布,经 The Good Investors 转录摘录核读):「I’m a big believer in scaling laws I’ll first say. In fact, if anything, the bet we placed in 2019 was on scaling laws and I stay on that. In other words, don’t bet against scaling laws. But at the same time, let’s also be grounded…」以及「these exponentials on scaling laws will become harder…But I would just still say…pre-training I think is not over, it continues.」同期 Microsoft Ignite(2024-11-19,经 TechCrunch 2024-11-20引用):「We are seeing the emergence of a new scaling law」;同场另一段(经 Marcus Fortune 稿转引):「the thing to remember at the end of the day these are not physical laws. There are just empirical observations that hold true just like Moore’s Law did for a long period of time and so therefore it’s actually good to have some skepticism some debate.」
Ilya Sutskever NeurIPS 演讲(2024-12-13,温哥华,2014 seq2seq 论文 Test of Time 获奖演说)。主笔经 The Verge 现场报道(Kylie Robison)亲核三句的各自位置:「Pre-training as we know it will unquestionably end, Sutskever said onstage.」「We’ve achieved peak data and there’ll be no more, according to Sutskever. ‘We have to deal with the data that we have. There’s only one internet.’」Reuters(经 Global Banking & Finance 转载核读)版:「While compute is growing, the data is not growing, because we have but one internet.」The Verge 另记其化石燃料比喻(图注:Ilya Sutskever calls data the “fossil fuel” of AI)与 hominid 类比——「just as evolution found a new scaling pattern for hominid brains, AI might similarly discover new approaches to scaling beyond how pre-training works today.」拼接引文警示:网传合并版「Pre-training as we know it will unquestionably end. The data is the fossil fuel of AI. We have but one internet.」是多句拼接,三句在演讲中位置不同;引用必须拆开标注来源(The Verge 现场记录 vs Reuters 稿 vs WSJ 转引)。
3. 第二幕(2024-12 → 2025-03):WSJ 稿、Reflections 与 DeepSeek 冲击
WSJ Orion 稿(2024-12 月下旬,Deepa Seetharaman、Berber Jin;标题双版本并存:URL/转载版《OpenAI’s Next Big AI Effort, GPT-5, Is Behind Schedule and Crazy Expensive》 vs Altman 引述版《The Next Great Leap in AI Is Behind Schedule and Crazy Expensive》,疑为改题或线上线下差异,引用并列注明)。要点(经 GIGAZINE 编译核读):GPT-5 研发 18+ 个月、至少两次大规模训练、2024-05 启动的大训发现「the training data was not as diverse as expected」、进展「not been enough progress to justify the enormous costs」。Altman 2024-12-22 推文反讽(逐字):「i think the wsj is the overall best us newspaper right now, but they published an article called ‘The Next Great Leap in AI Is Behind Schedule and Crazy Expensive’ many hours after we announced o3?!」——o3 公告见第五节。
命定派的高水位:Altman《Reflections》博客,2025-01-06:「We are now confident we know how to build AGI as we have traditionally understood it.」Amodei 达沃斯 2025-01-21(CNBC 场外;经 Ars Technica 2025-01-22转述标题即结论):AI 或于 2027 年后不久超越「almost all humans at almost everything」,并称 AGI 是营销术语、他偏好「a country of geniuses in a data center」。黄仁勋 CES 2025(Y2Doc 逐字转录,00:22:06):「And the scaling law continues. There are, in fact, two other scaling laws that have now emerged.」GTC 2025(Rev 逐字转录,约 40:18):「Well, this last year, this is where almost the entire world got it wrong. The computation requirement, the scaling law of AI is more resilient, and in fact, hyper accelerated. The amount of computation we need at this point…is easily 100 times more than we thought we needed this time last year.」
DeepSeek 冲击(2025-01-27):英伟达单日 -16.97%、市值蒸发约 $589B(史上最大单日市值损失,Investopedia 当日两篇)。「$5.5M 训练成本」叙事的出处是 DeepSeek-V3 技术报告 arXiv:2412.19437:「DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training」,按 $2/小时折算 $5.576M——但论文自己排除了前期成本,逐字:「Note that the aforementioned costs include only the official training of DeepSeek-V3, excluding the costs associated with prior research and ablation experiments on architectures, algorithms, or data.」SemiAnalysis《DeepSeek Debates》2025-01-31反驳(公开部分逐字):「the main headline being the ‘$6M’ dollar figure training cost of DeepSeek V3. This is wrong.」「We are confident that their GPU investments account for more than $500M」「the total server CapEx for DeepSeek is ~$1.6B」。Nadella 的回应(X 2025-01-27,经 Mashable India嵌入原文核到):「Jevons paradox strikes again! As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can’t get enough of.」SemiAnalysis 的框架句(对全篇有用):「the narrative has flipped from last month, when scaling laws were broken, we dispelled this myth, now algorithmic improvement is too fast and this too is somehow bad for Nvidia and GPUs.」
GPT-4.5 的系统卡自白(2025-02-27 发布,即 Orion 的出货形态):OpenAI 系统卡初版有一行被 Nathan Lambert(Interconnects,2025-02-28)逐字记录:「> GPT-4.5 is not a frontier model.」(正式博客版系统卡删去了此行——版本差异,引用必须注明出自初版。)Lambert 的定调句是中间派的代表文本:「The main contradiction to the claims that it isn’t a frontier model is that this is the biggest model the general public has ever gotten to test. Scaling to this size of model did NOT make a clear jump in capabilities we are measuring.」「Scaling language models is not dead. Still… We’ve entered the era where trade-offs among different types of scaling are real.」「…it was clear that scaling pretraining alone was not going to give us the same level of breakthroughs. Now we really know what Ilya saw.」
4. 第三幕(2025-08 → 2025-12):GPT-5、Marcus 宣判与泡沫之夏
GPT-5 发布与讣告复起(2025-08-07)。Altman 直播原话(经 Marcus 帖内转录):「With GPT-5 now it’s like talking to an expert —- a legitimate PhD level expert in anything any area you need on demand they can help you with whatever your goals are.」发布 48 小时内:3000 人请愿要求恢复 GPT-4o;Polymarket「8 月底最佳模型」盘口 OpenAI 一小时内从 75% 跌至 14%(同帖记录)。Gary Marcus 宣判帖(《GPT-5: Overdue, overhyped and underwhelming》,2025-08-09,逐字):「GPT-5 is barely better than last month’s flavor of the month (Grok 4); on some metrics (ARC-AGI-2) it’s actually worse. People had grown to expect miracles, but GPT-5 is just the latest incremental advance.」「That’s exactly what it means to hit a wall, and exactly the particular set of obstacles I described in my most notorious (and prescient) paper, in 2022.」「Pure scaling simply isn’t the path to AGI.」「Nobody with intellectual integrity should still believe that pure scaling will get us to AGI.」Marcus 的立场连续性值得标注:2022-03《Deep learning is hitting a wall》→ 2024-11「that’s just a fantasy」→ 2025-02 Fortune 稿「The fact is, pure scaling has not worked.」(Fortune 2025-02-19)→ 2025-08 宣判;其自我贴金「(and prescient)」是当事人修辞,不是第三方认证。
泡沫之夏(2025-08 下旬):Fortune 2025-08-24 专访(《’It’s almost tragic’》,链接)记 Altman 自承 GPT-5 发布「totally screwed up」;Altman 在记者晚宴谈泡沫(The Verge 转引):「When bubbles happen, smart people get overexcited about a kernel of truth,」;MIT NANDA「95%」报告病毒式传播(第六节详审);Eric Schmidt 2025-08-19 NYT 联署专栏转向:「it is uncertain how soon artificial general intelligence can be achieved」。
Ilya 的「研究时代」与部分收回(Dwarkesh Podcast,2025-11-25,官方逐字稿,时间戳 00:19:00 起):「Up until 2020, from 2012 to 2020, it was the age of research. Now, from 2020 to 2025, it was the age of scaling—maybe plus or minus, let’s add error bars to those years… But now the scale is so big. Is the belief really, ‘Oh, it’s so big, but if you had 100x more, everything would be so different?’ It would be different, for sure. But is the belief that if you just 100x the scale, everything would be transformed? I don’t think that’s true. So it’s back to the age of research again, just with big computers.」同段:「At some point though, pre-training will run out of data. The data is very clearly finite.」(00:36:38)「scaling sucked out all the air in the room. …we are in a world where there are more companies than ideas by quite a bit.」部分收回(2025-11-26~28 X,经 The Decoder 更新转述,推文原文未能抓取):Sutskever 澄清 current scaling technology 会继续带来改进、不会停滞,但「something important」仍会缺失。⚠️ 网传「The data is finite. There is only one internet. Pre-training as we have known it is over.」并非该访谈逐字句,是二手合并改写;逐字以转录为准。
年末反转:Altman 内部「code red」(2025-12-02 备忘录,The Information 首报,经二手转述「a critical time for ChatGPT」),随后密集发布 GPT-5.1(2025-11-12)、GPT-5.2(2025-12-11)、Google Gemini 3(2025-11)。
5. 2026 上半年:论战的当前落点
黄仁勋(Lex Fridman #494,2026-03,官方逐字稿,00:22:39 起):Lex 问「So are you still a believer in the scaling laws?」——「Yeah, yeah. Yeah, we have more scaling laws now.」(00:23:12)「And Ilya Sutskever said, ‘We’re out of data,’ or something like that. ‘Pre-training is over,’ or something like that. The industry panicked, you know, that this is the end of AI. And of course, that’s obviously not true.」(00:27:44)「And so the next scaling law is the agentic scaling law… you know, I have four scaling laws.」Amodei 达沃斯 2026-01(与 Hassabis 同场;经 TrendingTopics核读)坚持 2026/2027 诺奖级预测,并自曝收入链:「Our revenues grown 10x in the last three years from 0 to 100 million in 2023, 100 million to a billion in 2024 and 1 billion to 10 billion in 2025.」Hassabis 对照:本十年末前实现全人类认知能力的概率「50 percent」。Amodei(CBS 专访全文转录,2026-02-28):「AI is moving so fast…the amount of computation that goes into the models doubles every four months. We have never seen anything like this pace of innovation.」新代际节奏:Claude Opus 4.6(2026-02-05)、GPT-5.4(2026-03-05,105 万 token 上下文)——截至 2026-07-23,未检索到权威「盖棺」稿;7 月最近的信号是 TechCrunch 2026-07-09《Can AI answer the $3 trillion question?》(见第六节)。
6. 第二层裁决
时间线与原话是全部七层里最硬的材料——双方都留下了充足的逐字记录。两个文献学要点:(a) 「撞墙」是媒体框架,The Information 引爆原文只写「rate of improvement…slowing down」,wall 一词由 Bloomberg/Platformer/Marcus 注入;(b) 传输层病三种——Ilya 三合一拼接引文、WSJ 标题双版本、GPT-4.5 系统卡删行,引用任何一条都必须带版本注。至于双方原话本身:讣告派(Ilya「plateaued/unquestionably end」、Marcus 四连宣判)与命定派(Altman「no wall/confident」、Huang「intact/continues/more laws now」、Amodei「bet in favor」、Nadella「don’t bet against」)都在场,且两位最大押注者(Amodei、Nadella)都自承 scaling laws 是 empirical observations 不是 physical laws——名号跳是当事人自己先拆的。
三、数据墙层:「只有一个互联网」的会计学
Ilya 的「peak data」是讣告派最硬的物理论证。这一层审数据约束的定量账:存量多大、何时耗尽、缓冲有几层。
1. Villalobos 两版:一次公开的估计自我修正史
Villalobos et al.《Will we run out of data?》arXiv:2211.04325(v1 2022-10-26 → v2 2024-06-04;v2 为 ICML 2024 Position track,PMLR 页)。两版结论逐字对照:
- v1 摘要:「the stock of high-quality language data will be exhausted soon; likely before 2026. By contrast, the stock of low-quality language data and image data will be exhausted only much later; between 2030 and 2050 (for low-quality language) and between 2030 and 2060 (for images).」
- v2 摘要(主笔亲核):「if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained.」
官方博客的回述(Epoch AI 2024-06,逐字):「Our 2022 paper predicted that high-quality text data would be fully used by 2024, whereas our new results indicate that might not happen until 2028.」——注意 v1 摘要写「likely before 2026」、博客回述 v1 为「by 2024」,两个口径并存;修正来源是 Penedo et al. 2023(高质量存量上调约 5x)与 Muennighoff et al. 2023(多 epoch 有效存量再上调)。
v2 正文数字链(Table 1/Figure 3 逐字):索引网原始存量「around 510 trillion [95%: 130T, 2100T]」→「Quality-adjusted stock 100T [22T, 490T]」→「Repetition-adjusted stock 320T [65T, 1700T]」;耗尽年份「The median exhaustion year is 2028, and by 2032 exhaustion becomes very likely.」多 epoch 假设锚定 5x(「we reduce it to 5x」,脚注:典型 1–4 epochs);overtraining 敏感分析:「Llama 3 8B is overtrained by close to 100x, while Llama 3 70B is only overtrained [by ~10x]」;博客版补:「If models are overtrained by a modest factor of 5x, the stock of data will be fully used by 2027, but if they are overtrained by 100x, the stock of data will be fully used by 2025.」有效总存量口径(博客 Results 节逐字):「the total effective stock of human-generated public text data is on the order of 300 trillion tokens, with a 90% confidence interval of 100T to 1000T.」——不确定带宽一个数量级,这是「数据墙年份」类头条必须带上的误差棒。
2. Epoch《Can AI Scaling Continue Through 2030?》:四约束与「数据墙」原话
Epoch AI 2024-08-20(主笔亲核全文)。趋势锚(Introduction 逐字):「training compute expanding at a rate of approximately 4x per year」(对照:超过手机 2x/年、太阳能 1.5x/年、基因测序 3.3x/年的峰值增速)。四约束:「power availability, chip manufacturing capacity, data scarcity, and the ‘latency wall’」。数据小节逐字:「We estimate that the indexed web contains around 500 trillion tokens after deduplication, 30 times more data than the largest known training datasets.」「If the recent trend of 4x/year compute scaling continues, we would run into this ‘data wall’ for text data in about five years.」算上多模态与多 epoch:「we estimate the equivalent of 400 trillion to 20 quadrillion tokens available for training by 2030, allowing for 6e28 to 2e32 FLOP training runs.」底线(Bottom line 逐字):「training runs of around 2e29 FLOP are likely possible by 2030. This represents a significant increase in scale over current models, similar to the size difference between GPT-2 and GPT-4. The constraint likely to bind first is power, followed by the capacity to manufacture enough chips.」——在 Epoch 自己的排序里,数据不是第一约束,电力才是;「数据墙五年论」是有 4x/年趋势的明确条件句。(口径注:该文 Introduction 用「500T words」、正文用「500 trillion tokens」,原文自身不一致,引用以正文 tokens 为准。)
3. Model collapse 正反:Nature 的「不可避免」与 NeurIPS 的「累积可避免」
正方。Shumailov et al. 2024, Nature 631, 755–759(2024-07-24 在线;arXiv 原版 2305.17493《The Curse of Recursion》;2025-03-21 有 Author Correction,仅更正一处符号笔误)。Nature 版摘要逐字(主笔 Crossref 元数据亲核):「We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear. We refer to this effect as ‘model collapse’…」正文首节:「this process is inevitable, even for cases with almost ideal conditions for long-term learning, that is, no function estimation error.」以及数据墙直接相关句:「the use of LLMs at scale to publish content on the Internet will pollute the collection of data to train their successors: data about human interactions with LLMs will be increasingly valuable.」Definition 2.1:「In early model collapse, the model begins losing information about the tails of the distribution; in late model collapse, the model converges to a distribution that carries little resemblance to the original one, often with substantially reduced variance.」实验注意点:LLM 部分是 OPT-125m 在 wikitext2 上的微缩设定;且保留 10% 原始数据的设定下「only minor degradation of performance」——论文自己的数字里就含着「保留原始数据可缓解」的线索。版本注意:Nature 版摘要比 arXiv 版多了「indiscriminate」一词——限定词恰是两派分歧的七寸。
反方。Gerstgrasser et al. 2024《Is Model Collapse Inevitable?》arXiv:2404.01413(NeurIPS 2024)摘要逐字:「We confirm that replacing the original real data by each generation’s synthetic data does indeed tend towards model collapse, then demonstrate that accumulating the successive generations of synthetic data alongside the original real data avoids model collapse…」理论结果:替换时测试误差随代数线性增长;累积时「the test error has a finite upper bound independent of the number of iterations, meaning model collapse no longer occurs」(Theorem 2)。Discussion 逐字:「the ‘curse of recursion’ may not be as dire as had been portrayed – provided we accumulate synthetic data alongside real data, rather than replacing real data by synthetic data only.」保留条款(作者自报):VAE 实验中累积误差仍随代数增长(「albeit much more slowly」);且「’model collapse’ – as a term of art – has been used in various ways by various researchers; so care is required in comparing claims across articles.」Dohmatob et al. 2024 arXiv:2402.07712 在高维回归设定给出坍缩的解析刻画与缓解策略;其团队后续 arXiv:2410.04840《Strong Model Collapse》 称监督回归设定下「as little as 1% of the total training dataset」的合成数据也可致坍缩——设定边界必须随文带上。
裁决小注:「合成数据必然毒化」与「合成数据随便用」都不立。一手文献的实际交集是:indiscriminate/replace 用法坍缩,accumulate/精心配比可用——而这正是业界的实际用法(见下条 phi-4)。
4. 合成数据与数据筛选的工程实证
phi 系列。Gunasekar et al. 2023《Textbooks Are All You Need》arXiv:2306.11644:phi-1(1.3B)用「textbook quality」数据(6B web tokens + 1B GPT-3.5 合成)达到 HumanEval 50.6%。Abdin et al. 2024《phi-4》arXiv:2412.08905:「Synthetic data constitutes the bulk of the training data for phi-4」,预训练约 10T tokens,合成数据占混合配比 40%、web 与 web 改写各 15%、代码 20%、定向获取 10%;并声称「phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities」。但论文自己的消融保留条款(§4 逐字)是双向审查的好材料:「Models trained only with synthetic data underperformed on the knowledge-heavy benchmarks and demonstrated increased hallucinations.」「Prioritizing more epochs over our synthetic data led to better performance with respect to adding fresh web tokens.」——纯合成不行、多 epoch 优于堆新 web:这两句同时给讣告派(合成有极限)和命定派(数据墙有多 epoch 缓冲)递刀。
数据筛选。DataComp(Gadre et al. 2023)arXiv:2304.14108(⚠️ 任务书曾误植为 2304.14177):DataComp-1B 以同训练程序同算力把 CLIP ViT-L/14 zero-shot ImageNet 做到 79.2%、超 OpenAI CLIP 3.7 个百分点——数据质量本身是收益轴。DCLM(Li et al. 2024)arXiv:2406.11794:「we provide a standardized corpus of 240T tokens extracted from Common Crawl」,DCLM-Baseline 用 2.6T tokens 从零训出 7B 模型 MMLU 64%、比前开源数据 SOTA 高 6.6 点且少用 40% 算力。Goyal et al. 2024《Scaling Laws for Data Filtering》arXiv:2404.07177(CVPR 2024):「the limited high-quality data rapidly loses its utility when repeated, eventually requiring the inclusion of ‘unseen’ but ‘lower-quality’ data. Our key message is that data curation cannot be agnostic of the total compute that a model will be trained for.」
多 epoch 实证。Muennighoff et al. 2023《Scaling Data-Constrained Language Models》arXiv:2305.16264(NeurIPS 2023;v5 2025-06-28)摘要逐字(主笔亲核):「training with up to 4 epochs of repeated data yields negligible changes to loss compared to having unique data. However, with more repetition, the value of adding compute eventually decays to zero.」§6 定量:8.7B 模型训 4 epoch(44B unique tokens)比单 epoch(178B unique tokens)验证 loss 仅高 0.5%;重复「半衰期」R_D≈15——「Meaningful gains from repeating data can be made up to around 16 epochs (R_D) beyond which returns diminish extremely fast.」
5. 业界实际数据量与 2025–2026 更新
前沿模型 token 预算公开链(全部官方一手):Llama 2→3(Llama 3 herd 论文 arXiv:2407.21783 §1):「about 15T multilingual tokens, compared to 1.8T tokens for Llama 2」(旗舰 405B 用 15.6T);DeepSeek-V3 arXiv:2412.19437:「14.8 trillion diverse and high-quality tokens」;Qwen3 arXiv:2505.09388:「a total of 36 trillion tokens」;Llama 4(Meta 官方 HF 模型卡):「Llama 4 Scout was pretrained on ~40 trillion tokens and Llama 4 Maverick was pretrained on ~22 trillion tokens」——网传「15T→40T 统一数字」不实,两款口径不同。
Epoch 2025《AI in 2030》(PDF 全文,Google DeepMind 委托)Executive summary 逐字:「We show how general-purpose publicly available text data plausibly could be exhausted before 2027. Nevertheless, AI developers are unlikely to run out of data for large-scale training. This is due to two outstanding sources: synthetic data (particularly for reasoning training), and multi-modal data.」数据章节:「Datasets for general-purpose AI training recently grew at 2.7x per year…」——Epoch 自己的最新立场:公共文本可见顶,但「耗尽」不等于「无数据可用」。
版权线(NYT v. OpenAI,SDNY 1:23-cv-11195):2023-12-27 起诉;2025-03 撤案动议核心部分被驳回进入实体审理;2025-05-13 Magistrate Judge Wang 签发保全令要求保留全部输出日志;2026-01-05/06 Judge Stein 维持裁定、强制交付 2000 万条去标识化对话(ABA Journal 2026-01-08);OpenAI 官网(openai.com/new-york-times)逐字:「We have been actively fighting the Times’ demand that we turn over 20 million of your private ChatGPT conversations.」——截至 2026-07 未产生终局侵权/合理使用判决;对数据墙的约束目前体现在合规成本与数据获取结构,而非法律禁运。
6. 第三层裁决
数据约束真:公共人类文本存量有限(300T tokens 级、带宽一个数量级)、随 4x/年算力趋势在 2026–2032 窗口见顶,这是讣告派最实的物理锚。但作为「撞墙终审」未立:(a) 估计本身已被作者公开修正过一次(v1→v2 后移 4 年);(b) 三层缓冲各有实证——多 epoch(4 epoch 近似无损、16 epoch 半衰)、合成数据(phi-4 40% 配比且超越 teacher、collapse 之争以「累积可用」收场)、多模态与筛选(DCLM 240T 池);(c) Epoch 2025 自判「unlikely to run out of data」;(d) 业界 token 预算仍在翻倍(15T→36T→40T)。「只有一个互联网」作为修辞成立,作为终审不成立——且 Ilya 本人在 2025-11 也退到「会继续改进、但缺某个重要东西」。
四、代际证据层:到底放缓没有
讣告派的最硬证据是单品(Orion/GPT-4.5 不及预期),命定派的最硬证据是曲线(2025–2026 代际节奏与基准轨迹)。这一层把两边可核的数字摆到同一张桌上。
1. GPT-4.5/Orion:讣告派最硬的单品证据
OpenAI 官方博文,2025-02-27(Wayback 2025-03-01 快照,主笔亲核):「We’re releasing a research preview of GPT‑4.5—our largest and best model for chat yet.」(⚠️ 媒体广泛引用的「largest and most knowledgeable model yet」并非博文首句原文——The Verge 称该措辞出自发布前泄露的文档;引用时应区分口径。)博文自己的标度叙事:「GPT‑4.5 is an example of scaling unsupervised learning by scaling up compute and data…」,结论段:「With every new order of magnitude of compute comes novel capabilities.」但同一篇博文的 API 段却在自我降温(主笔亲核):「GPT‑4.5 is a very large and compute-intensive model, making it more expensive than and not a replacement for GPT‑4o. Because of this, we’re evaluating whether to continue serving it in the API long-term…」官方定价页(Wayback 2025-03-03 快照):GPT-4.5 输入 $75/输出 $150 每百万 token——输入 30 倍、输出 15 倍于 GPT-4o($2.50/$10)。官方基准表:GPQA 71.4%(o3-mini high 79.7%)、AIME’24 36.7%(o3-mini high 87.3%)、SWE-bench Verified 38.0%(o3-mini high 61.0%)——自家表里就被自家推理模型全面压制。
媒体评价(逐字)。The Verge 2025-02-27 标题即「OpenAI announces GPT-4.5, warns it’s not a frontier AI model」,正文记录系统卡先写后删事件:「’GPT-4.5 is not a frontier model…It does not introduce 7 net-new frontier capabilities compared to previous reasoning releases, and its performance is below that of o1, o3-mini, and deep research on most preparedness evaluations.’ OpenAI has since removed this mention from an updated version of the document.」Altman 同日 X 帖(经该文转引):「giant, expensive model」「won’t crush benchmarks」。Ars Technica 2025-02-28 标题「’It’s a lemon’…arrives to mixed reviews」,副题「marginal gains in capability and poor coding performance despite 30x the cost」;Karpathy 的 X 帖(Ars 转引)是中间派样本:「Everything is a little bit better and it’s awesome, but also not exactly in ways that are trivial to point to.」
承重判断:Orion 不及预期是真——被 The Information/Bloomberg/WSJ 三路独立报道、被 OpenAI 自己的系统卡初版与定价(30 倍于 4o)双重确认。这是讣告派全部论据里最实的一件。但它的射程是一件产品,不是一条曲线。
2. GPT-5:官方数字与舆论裂口
OpenAI 官方博文,2025-08-07(主笔亲核):「It sets a new state of the art across math (94.6% on AIME 2025 without tools), real-world coding (74.9% on SWE-bench Verified, 88% on Aider Polyglot), multimodal understanding (84.2% on MMMU), and health (46.2% on HealthBench Hard)」,GPT-5 pro 扩展推理「sets a new SOTA on GPQA, scoring 88.4% without tools」(脚注:SWE-bench 用 n=477 固定子集;带工具 AIME 不与无工具直接可比)。定价 $1.25/$10——比 GPT-4.5 便宜 60 倍。
Altman 吹风会原话(The Verge 2025-08-07,Alex Heath 现场,比 Marcus 帖内转录更可靠):「GPT-3 sort of felt like talking to a high school student… GPT-4 felt like you’re talking to a college student. GPT-5 is the first time that it really feels like talking to a PhD-level expert.」「This is clearly a model that is generally intelligent.」——但同场承认缺口(这句常被漏引):「This is not a model that continuously learns as it’s deployed from the new things it finds, which is something that, to me, feels like it should be part of AGI.」
发布混乱三件事(全部一手转引):①「chart crime」——Futurism 2025-08-08:「Somehow, the bar for GPT-5’s score of 52.8 percent accuracy is nearly twice as tall as the bar for a score of 69.1 percent for the o3 model.」Altman 自称「mega chart screwup」(TechCrunch 2025-08-08,同文注明博客正式版图表是对的);②路由器故障——Altman Reddit AMA(TechCrunch 转引):「Yesterday, we had a sev and the autoswitcher was out of commission for a chunk of the day, and the result was GPT-5 seemed way dumber.」③GPT-4o 下架与恢复——Ars 2025-08-13:「GPT-4o has returned to ChatGPT following intense user backlash」,Altman 承认低估「how much some of the things that people like in GPT-4o matter to them」。代表性评测 Simon Willison 2025-08-07 是平衡样本:「It’s my new favorite model…It doesn’t feel like a dramatic leap ahead from other LLMs but it exudes competence—it rarely messes up, and frequently impresses me. The pricing is aggressively competitive.」
承重判断:GPT-5 的官方数字是真进步(AIME 无工具 94.6%、SWE 74.9% 创 SOTA),但发布事件本身成了讣告的燃料——「PhD-level expert」修辞 vs「seemed way dumber」首日体验 vs 4o 恢复事件,舆论裂口是真实的。后续轨迹:GPT-5.1(2025-11-12,主打「更温暖」回应性格反弹)、GPT-5.1-Codex-Max(SWE 77.9% xhigh,宣称内部观察到「连续工作超过 24 小时」任务、「95% 的 OpenAI 内部工程师每周使用 Codex」)、GPT-5.2(2025-12-11,「code red」后提前发布——该叙事经二手汇总转引 Reuters/Axios,标注待核)。2026 年翻案叙述代表 Pouladian 2026-03-06:「the scaling laws skeptics are wrong」「Scaling didn’t hit a wall. It found a new dimension.」
3. Claude 链:另一条爬坡曲线(SWE-bench 33.4% → 93.9%)
Anthropic 官方口径的代际链(全部为官方博文,数字口径差异随注):3.5 Sonnet(新),2024-10-22:「On coding, it improves performance on SWE-bench Verified from 33.4% to 49.0%, scoring higher than all publicly available models…」→ 3.7 Sonnet,2025-02-24:「the first hybrid reasoning model on the market」,SWE 62.3%/63.7%(n=489)、高算力 70.3% → Opus 4 / Sonnet 4,2025-05-22:SWE 72.5%/72.7%、高算力 79.4%/80.2% → Sonnet 4.5,2025-09-29:SWE 77.2%、OSWorld 61.4%(「Just four months ago, Sonnet 4 held the lead at 42.2%」)、「maintaining focus for more than 30 hours on complex, multi-step tasks」→ Opus 4.5,2025-11-24(主笔亲核):「the best model in the world for coding, agents, and computer use」,$5/$25(Opus 级价格脚踝斩),内部招聘级考试「Within our prescribed 2-hour time limit, Claude Opus 4.5 scored higher than any human candidate ever」——脚注必须带:该成绩用了 parallel test-time compute,且「Without a time limit, the model…matched the best-ever human candidate」;SWE-bench 80.9%(⚠️ 该数字以图表呈现,经第三方一致转引,标注待 system card 复核)→ Opus 4.6,2026-02-05:SWE 80.8%(第三方转引官方表格)——比 4.5 的 80.9% 略低,代际增益非单调的活样本 → Opus 4.7,2026-04-16:SWE 87.6%(第三方多家一致转引)→ Mythos Preview / Project Glasswing,2026-04-07(主笔亲核):「AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.」「Mythos Preview has already found thousands of high-severity vulnerabilities, including some in every major operating system and web browser.」——OpenBSD 27 年远程崩溃漏洞、FFmpeg 16 年漏洞(自动化测试命中 500 万次未发现)、Linux 内核提权链;SWE-bench Verified 93.9% vs Opus 4.6 80.8%(官方注明做了记忆化筛查且排除后优势不变);HLE 有工具 64.7% vs 53.1%;「We do not plan to make Claude Mythos Preview generally available」。后续:Mythos 5 / Fable 5(2026-06-09,审查准入)、07-01 官方宣布出口管制解除。
4. Gemini 翻转与 Llama 4 反面教材
Google:Gemini 2.5 Pro,2025-03-25:「debuts at #1 on LMArena by a significant margin」,HLE 18.8% → Gemini 3 Pro,2025-11-18(主笔亲核):「It tops the LMArena Leaderboard with a breakthrough score of 1501 Elo.」HLE 37.5%(无工具)、GPQA Diamond 91.9%、SWE-bench 76.2%;Deep Think:HLE 41.0%、ARC-AGI-2 45.1%(ARC Prize Verified)→ Gemini 3.1 Pro(preview),2026-02-19:「On ARC-AGI-2…3.1 Pro achieved a verified score of 77.1%. This is more than double the reasoning performance of 3 Pro.」——ARC-AGI-2 从 o3-preview 的约 4%(2025-03)到 77.1%(2026-02)用了 11 个月,这是「新基准被快速攻陷」命题的最硬单例。
Meta(讣告派的另一件真证据):Llama 4(2025-04-05,Wayback 快照)官方自称 Maverick「with an experimental chat version scoring ELO of 1417 on LMArena」——「experimental chat version」字样就在官方博文里。LMArena 官方声明(2025-04-08,经 THE DECODER 逐字互证):「Meta’s interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that ‘Llama-4-Maverick-03-26-Experimental’ was a customized model to optimize for human preference.」公开版真实排名约第 32 位。Behemoth 两度推迟(WSJ 2025-05-15 经 GIGAZINE 转述),2026 年状态「effectively shelved」(二手);Meta 转向 MSL/Avocado/Muse Spark 线(2026-04-08 首个闭源前沿模型;内部备忘录称效率 10× Maverick——均二手,标注待核)。承重判断:开源权重阵营的代际放缓是真的(Llama 3 405B→Llama 4 争议→Behemoth 搁置),但它证明的是 Meta 一家的执行问题,不是 scaling 的物理墙。
5. 基准饱和与换基准军备竞赛
饱和是真。MMLU-Pro arXiv:2406.01574 设立动机逐字:「as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities.」HLE arXiv:2501.14249:「LLMs now achieve over 90% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities.」Stanford HAI AI Index 2025:「scores rose by 18.8, 48.9, and 67.3 percentage points on MMMU, GPQA, and SWE-bench, respectively」——一年内;「the score difference between the top and 10th-ranked models fell from 11.9% to 5.4% in a year, and the top two are now separated by just 0.7%」;「training compute doubles every five months, datasets every eight」;GPT-3.5 级推理成本 2022-11→2024-10 降 280 倍。
但换基准的攻陷速度同样真。FrontierMath(Epoch AI 2024-11):设立时「Current AI systems solve less than 2%」,Tao:「I think they will resist AIs for several years at least.」Gowers:「Getting even one question right would be well beyond what we can do now」——轨迹:<2%(2024-11)→ o3 25.2%(2024-12;⚠️ 污染争议:arXiv:2508.00459 脚注指 OpenAI 资助 FrontierMath 且曾接触几乎全部题目,Epoch 仅保留 50 题独立 holdout)→ GPT-5.2 Thinking 40.3%(官方)→ GPT-5.4 Thinking 47.6% / Pro 50.0%(Tier 1–3,GPT-5.5 官方对比表)→ GPT-5.5 51.7%。HLE:<10%(2025-01 论文基线)→ 18.8%(Gemini 2.5)→ 37.5%(Gemini 3)→ 64.7%(Mythos Preview,有工具;官方自注低 effort 表现或含记忆化)——18 个月从个位数到约 65%。
SWE-bench Verified 轨迹(跨口径警示:OpenAI 用 n=477、Anthropic 多用全 500 题、scaffold 各异,不宜严格同口径比较):33.4%(2024-06 Claude 3.5)→ 49.0%(2024-10)→ 62–64%(2025-02/03,3.7/Gemini 2.5)→ 72–74%(2025-05/08,Claude 4/GPT-5)→ 76–81%(2025-11/12,Gemini 3/Opus 4.5/GPT-5.2)→ 87.6%(2026-04 Opus 4.7)→ 93.9%(Mythos Preview)。水分警示:METR 2026-03《Many SWE-bench-Passing PRs Would Not Be Merged into Main》:「roughly half of test-passing SWE-bench Verified PRs written by recent AI agents would not be merged into main by repo maintainers」,自动评分平均虚高约 24 个百分点。
6. METR 时间视界:最像「没撞墙」的单条证据
METR 2025-03-19(主笔亲核):「The length of tasks (measured by how long they take human professionals) that generalist frontier model agents can complete autonomously with 50% reliability has been doubling approximately every 7 months for the last 6 years.」现状刻画:「current models have almost 100% success rate on tasks taking humans less than 4 minutes, but succeed <10% of the time on tasks taking more than around 4 hours」;置信表述:「we’re fairly confident that the overall trend is roughly correct, at around 1-4 doublings per year」;SWE-bench Verified 独立子集「shows an even faster doubling time, of under 3 months」。追踪页(metr.org/time-horizons)逐月收录新模型:GPT-5 约 2 小时 17 分(2025-08)、Opus 4.6 约 14.5h(95% CI 6–98h,风险报告自评「extremely noisy」)、Mythos Preview 点估计约 17h 并挂出「Measurements above 16 hrs are unreliable with our current task suite」——且 2026-03-03 官方自我勘误「Corrected a regularization mistake that affected our measurements」,2026-02/03 风险报告自承「Recent models have been getting closer to saturating this underlying task suite, which makes the estimated time horizon less constrained by the data and more sensitive to analysis decisions.」批评面:「Task selection is the entire result.」(EA Forum 2026-07);「约 4 个月/约 3 个月翻倍加速」是第三方解读非 METR 原文结论。承重判断:METR 曲线是「能力仍在指数增长」命题的最干净单条证据,但其作者自己标注了饱和与测量敏感性——引用它做「没撞墙」终审同样越界。
7. 第四层裁决
双向各得一分,且分数都不小。讣告派得分:Orion/GPT-4.5 不及预期被官方文件确认;Llama 4/Behemoth 证明至少一家的预训练代际路线翻车;MMLU 类老基准饱和到失去区分度(top2 只差 0.7%)。命定派得分:2024-10 到 2026-04,SWE-bench 从 33.4% 爬到 93.9%、HLE 从个位数到约 65%、ARC-AGI-2 从约 4% 到 77.1%、METR 时间视界从约 1 小时到 10 小时+——「撞墙后 18 个月」恰是公开可核数字增长最快的 18 个月。两边都越界的地方:讣告派拿单品宣判曲线(GPT-4.5 一件产品→「scaling 已死」),命定派拿曲线洗白单品(新基准攻陷→Orion 事件不存在)。而 METR 自己挂出的「above 16 hrs unreliable」与 Anthropic 的「记忆化筛查」脚注提醒:连最硬的反讣告证据都带着自报的测量天花板。
五、新轴层:推理时计算——墙叙事为什么死了,以及它自己的天花板
「撞墙」论战在 2025 年事实上熄火,主因不是谁辩赢了,而是行业换了一条轴并且拿出了数字。这一层审:新轴兑现了多少、它自己的天花板在哪。
1. o1:新轴的出生证(2024-09-12)
OpenAI《Learning to Reason with LLMs》(官网 / Wayback 2024-09-13 快照,主笔亲核):「We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). The constraints on scaling this approach differ substantially from those of LLM pretraining, and we are continuing to investigate them.」图注:「o1 performance smoothly improves with both train-time and test-time compute」(注意:「对数-对数坐标」只能从图本身看到,正文无 “log-log” 字样)。AIME 数字两个口径必须并列:正文「GPT-4o only solved on average 12% (1.8/15)…o1 averaged 74% (11.1/15) with a single sample」;附录 A 表格口径(「74.4%」的真正出处):AIME 2024 pass@1 — gpt-4o 9.3 / o1-preview 44.6 / o1 74.4;cons@64 83.3;GPQA Diamond pass@1 77.3;CodeForces Elo 1,673(89 百分位)。(o1 System Card arXiv:2412.16720 是安全文档,不含 AIME 跑分——网传「系统卡载 AIME 74.4%」系误植;卡内能力数字仅有 MLE-bench「outperform GPT-4o by at least 6%」。)Noam Brown 的定量直觉(TED AI 2024-10,VentureBeat 2024-10-23):「We’re no longer constrained to just scaling up the system one training. Now we can scale up the system two thinking as well, and the beautiful thing about scaling up in this direction is that it’s largely untapped.」
2. o3 与 ARC-AGI:新轴的烟花与账单(2024-12-20)
ARC Prize 官方博客(François Chollet 署名,2024-12-20;现行页含两次定价修订注,主笔亲核):「scored a breakthrough 75.7% on the Semi-Private Evaluation set at our stated public leaderboard $10k compute limit. A high-compute (172x) o3 configuration scored 87.5%.」成本三个版本必须并列:2024-12 原始版「o3 requires $17-20 per task in the low-compute mode」→ 现行页(经 2025-03、2025-12 两次改价)Semi-Private 高效率 $26/任务、低效率 $4,560/任务(5.7B tokens);人类对照「you could pay a human to solve ARC-AGI tasks for roughly $5 per task (we know, we did that)」。训练注记:「OpenAI shared they trained the o3 we tested on 75% of the Public Training set.」版本警示(页顶 2025-04-16 更新):正式版 o3 与本帖所测 o3-preview 不是同一模型。Chollet 本人三句双向(逐字):「This is not merely incremental improvement, but a genuine breakthrough, marking a qualitative shift in AI capabilities compared to the prior limitations of LLMs.」「o3’s improvement over the GPT series proves that architecture is everything. You couldn’t throw more compute at GPT-4 and get these results. Simply scaling up the things we were doing from 2019 to 2023 — take the same architecture, train a bigger version on more data — is not enough. Further progress is about new ideas.」「Passing ARC-AGI does not equate to achieving AGI, and, as a matter of fact, I don’t think o3 is AGI yet.」——注意第二句同时是讣告派弹药(2019–2023 那套 scaling 不够了)与命定派弹药(新范式真突破):Chollet 本人把它写在同一段里。
3. DeepSeek-R1:新轴的开源复现与边界自白
DeepSeek-R1,Nature 645, 633–638 (2025)(arXiv 版 2501.12948,v2 2026-01-04 与 Nature 版对齐)。正文逐字:「the average pass@1 score on AIME 2024 shows a marked increase, jumping from an initial value of 15.6% to 77.9%. Also, by using the self-consistency decoding, the performance of the model can be further improved, achieving an accuracy of 86.7%.」(注意:这三个数字都是 R1-Zero,不是 R1 本体。)纯 RL 主张:「we bypass the conventional supervised fine-tuning (SFT) phase before RL training」;边界自白(Conclusion/limitation 逐字):「Consequently, for complex tasks that cannot be effectively evaluated by a reliable reward model, scaling up pure RL methods remains an open challenge.」以及「for tasks that cannot obtain a reliable signal, DeepSeek-R1 uses human annotation to create supervised data and only conducts RL for hundreds of steps.」
4. 学界的定量:4x、14x 的条件句与 sigmoid 天花板
Snell et al. 2024 arXiv:2408.03314(主笔摘要亲核):「we can improve the efficiency of test-time compute scaling by more than 4x compared to a best-of-N baseline. Additionally, in a FLOPs-matched evaluation, we find that on problems where a smaller base model attains somewhat non-trivial success rates, test-time compute can be used to outperform a 14x larger model.」——两个数字都带条件:「4x」是 revisions+search 并用时;「14x」限定「小模型已有非平凡成功率的题目」。同一篇的反向句(常被漏引,Introduction):「That said, with the most challenging questions, we observe very little benefits from scaling up test-time compute…demonstrating that current approaches to scaling test-time compute may not be 1-to-1 exchangeable with scaling pretraining.」(ICLR 2025 正式版标题措辞与 arXiv 版不同,引用注明版本。)Wu et al. 2024《Inference Scaling Laws》arXiv:2408.00724(⚠️ 任务书曾误植为 2406.15762):「Llemma-7B…consistently outperforms the Llemma-34B model across all tested inference strategies on the MATH benchmark.」Akyürek et al. 2024 arXiv:2411.07279:TTT 在 ARC 上「up to 6x improvement in accuracy compared to base fine-tuned models」(⚠️ 任务书误记为 8x),8B 模型 53%、集成后 61.9%「matching the average human score」。
天花板证据。Meta《The Art of Scaling RL Compute》arXiv:2510.13786(主笔摘要亲核):「Unlike pre-training, which typically uses power-law to fit predictive curves, we model pass rate versus log(compute) with a sigmoidal function…we found the sigmoidal fit to be much more robust and stable compared to power law empirically.」「RL Performance Ceilings are Not Universal」——400,000 GPU 小时的实测结论:不同配方撞到不同渐近线,工程细节主要调节算力效率「without materially shifting the asymptote」。《The Art of Scaling Test-Time Compute》arXiv:2512.02008(8 个开源模型、300 亿+ tokens):「(1) no single TTS strategy universally dominates…(3) for a given model type, the optimal TTS performance scales monotonically with compute budget.」《Overthinking in LLM Test-Time Compute Scaling》arXiv:2604.10739(2026-04):「marginal returns diminish substantially at higher budgets and that models exhibit ‘overthinking’, where extended reasoning is associated with abandoning previously correct answers.」Apple《The Illusion of Thinking》arXiv:2506.06941:「LRMs face a complete accuracy collapse beyond certain complexities…their reasoning effort increases with problem complexity up to a point, then declines despite having remaining token budget.」
5. 各家定调与 ARC-AGI-2 的另类压力测试
OpenAI(o3/o4-mini 发布博文,2025-04-16,Wayback 亲核):「we’ve observed that large-scale reinforcement learning exhibits the same ‘more compute = better performance’ trend observed in GPT‑series pretraining. By retracing the scaling path—this time in RL—we’ve pushed an additional order of magnitude in both training compute and inference-time reasoning, yet still see clear performance gains…」Anthropic:3.7 Sonnet「the first hybrid reasoning model on the market」;Claude 4 附录自曝对照:AIME 无扩展思考仅 Opus 4 33.9%/Sonnet 4 33.1%(思考轴贡献的定量自证)。Google:Gemini 2.5「thinking models」。NVIDIA(Q4 FY2025 财报新闻稿,2025-02-26):「Demand for Blackwell is amazing as reasoning AI adds another scaling law — increasing compute for training makes models smarter and increasing compute for long thinking makes the answer smarter.」
ARC-AGI-2(ARC Prize 2025-03-24):「Pure LLMs score 0% on ARC-AGI-2, and public AI reasoning systems achieve only single-digit percentage scores. In contrast, every task in ARC-AGI-2 has been solved by at least 2 humans in under 2 attempts.」起步表:o3-preview-low 约 4%($200/任务)、o1-pro 1%、R1 0.3%、GPT-4.5 0.0%;人类平均 60%($17/任务)。ARC Prize 2025 结果(2025-12-05,Mike Knoop 署名):最高已验证商用模型 Claude Opus 4.5 (Thinking, 64k) 37.6%($2.20/任务);最高 refinement 方案 Gemini 3 Pro + Poetiq 54%($30/任务);Knoop 判词:「The invention and scale up of chain-of-thought synthesis rivals the invention and scale up of transformers. And yet we are still very early.」并自曝测量现实:「According to OpenAI, only ~10% of ChatGPT free users have ever used ‘thinking’ mode.」再到 2026-02 Gemini 3.1 Pro 的 77.1%(见第四节)——九个月从 4% 到 77%。
6. 第五层裁决
新轴真兑现:o1 的「differ substantially」被随后 21 个月证实——RL/test-time 轴贡献了 2025–2026 几乎所有头条增益(AIME 74.4%→99.5%、ARC-AGI 75.7%→ARC-AGI-2 77.1%、SWE 38%→93.9%),OpenAI 官方声称 RL 轴又走完「一个数量级」且「more compute = better performance」趋势同形。这是「撞墙」叙事事实上死亡的主因。但三条自限必须带上:(a) 学界已测到新轴自己的天花板形状——sigmoid(ScaleRL)、边际递减与「放弃正确答案」(Overthinking)、无通用最优策略(2512.02008)、高复杂度崩坍(Apple);(b) 收益集中在有可验证奖励的域——R1 自己的 limitation 写明纯 RL 对无可靠奖励任务「remains an open challenge」;(c) 「换轴」本身是对讣告派核心观察的确认——如果预训练单轴没有放缓,行业不需要换轴。Chollet 那句「architecture is everything」把两边的分都记在了同一段里:2019–2023 的 scaling 确实不够了(讣告派得分),新想法确实来了(命定派得分)。
六、经济账层:钱往哪走,账合不合
「撞墙」之争有一条不依赖任何基准分的裁判:资本开支。如果行业真相信预训练已死,钱应该撤出;如果账真合不上,钱应该缩水。两个检验的结果分裂。
1. 押注侧:撞墙论后 18 个月,capex 翻倍
一手财报链(全部为官方 IR/SEC 文件):Microsoft——Brad Smith 博客 2025-01-03:「In FY 2025, Microsoft is on track to invest approximately $80 billion to build out AI-enabled datacenters…」→ FY26 Q2(2026-01-28)单季 capex $37.5B → FY26 Q3 电话会(2026-04-29):「For calendar year 2026, we expect to invest roughly $190 billion in capital expenditures which includes approximately $25 billion from the impact of higher component pricing.」Nadella 同场:「Our AI business surpassed $37 billion ARR, up 123%.」Alphabet——2025 指引从 $75B(Q4 2024 电话会)一路上调至 $91–93B;FY2025 年报(SEC)实际 $91.4B;2026 指引 $175–185B;Q1 2026 实际 $35.7B(10-Q)。Meta——Q4 2025 财报:FY2025 capex $72.2B,2026 指引「$115-135 billion…driven by increased investment to support our Meta Superintelligence Labs efforts」→ Q1 2026上调至 $125–145B。Amazon——Q4 2025 财报(GeekWire 全文):Jassy「we expect to invest about $200 billion in capital expenditures across Amazon in 2026」,盘后 -10%。四大 2026 合计口径接近 $700B(CNBC 2026-02 汇总)。——2024-11「撞墙」引爆时,这四家年 capex 合计约 $250–300B 量级;论战 18 个月,押注 scaling 的钱翻了一倍多。
2. 收入侧:增长真,缺口也真
OpenAI:2024 收入约 $3.7B、亏损约 $5B(NYT 2024-09 经 PYMNTS转引)→ 2025 收入约 $13B、经营亏损约 $21B(Fortune 2026-06-16 泄露财务文件:「In 2024, the company spent $2.37 to generate every $1 in revenue. In 2025, that ratio declined to $1.60.」)→ ARR 口径 2025 年末超 $20B(CFO Friar 官方博客,经 Reuters 转引)→ 2026-01 末年化 $25B(The Information,经 Dealroom转引)。(口径提醒:「$13B 实际收入」与「$20B+ 年末 ARR」并存,引用勿混。)算力承诺:2025Q4 累计签约超 $1.4T(AFP 2025-11-07);HSBC 测算其到 2030 年仍有 $207B 资金缺口(Fortune 2025-11-26)。Altman 面对质疑的原话(BG² 播客,TechCrunch 2025-11-02):「First of all, we’re doing well more revenue than that. Second of all, Brad, if you want to sell your shares, I’ll find you a buyer. I just — enough.」收缩信号(2026):OpenAI 将算力支出计划削减至四年 $600B(DCD 2026-02-23 转引 The Information:「under half of its original $1.4 trillion commitment」);Friar 内部质疑 $600B 规模与 IPO 节奏(2026-04,经 Notebookcheck 转引 Reuters/The Information);「backstop」风波(2025-11-05→07)三方原话全链:Friar(WSJ Tech Live,经多家转引)「…maybe even governmental partners…can really drop the cost of the financing」,被追问是否指联邦 backstop 答「Exactly.」→ Friar 当日 LinkedIn 撤回(TechCrunch 2025-11-06):「OpenAI is not seeking a government backstop for our infrastructure commitments. I used the word ‘backstop’ and it muddied the point.」→ 白宫 AI 沙皇 David Sacks:「There will be no federal bailout for AI. The U.S. has at least 5 major frontier model companies. If one fails, others will take its place.」Anthropic:收入 0→$100M(2023)→$1B(2024)→$10B(2025,Amodei 达沃斯 2026 自曝);内部预测 2028 年收入 $70B(The Information 2025-11)。token 用量链(Alphabet 官方):9.7T/月(2024-05 I/O)→ 480T/月(2025-05 I/O)→「we have doubled that number now processing over 980 trillion monthly tokens」(Q2 2025 电话会,2025-07-23,Pichai,主笔亲核 IR 页)→ 约 1.3Q/月(2025-10,纪要类来源)→ 3.2 quadrillion/月(2026-05-20 I/O 2026,经 Economic Times 等转引)——两年约 330 倍。
3. 缺口叙事及其检验
Cahn 的缺口会计。David Cahn(Sequoia)《AI’s $600B Question》2024-06-20(主笔亲核全文):算法逐字——「take Nvidia’s run-rate revenue forecast and multiply it by 2x to reflect the total cost of AI data centers…Then you multiply by 2x again, to reflect a 50% gross margin for the end-user of the GPU.」(脚注:NVIDIA 自家 2023-10 analyst day 第 14 页有同一口径。)「The $125B hole is now a $500B hole」;著名的「妄想」段:「That delusion says that we’re all going to get rich quick, because AGI is coming tomorrow, and we all need to stockpile the only valuable resource, which is GPUs.」结尾平衡句必须带:「In reality, the road ahead is going to be a long one. It will have ups and downs. But almost certainly it will be worthwhile.」后续立场两次更新:《AI in 2026: A Tale of Two AIs》2025-12-03:「2026 will be a year of delays, first in data center buildouts…and second, in the AGI timeline.」收入缺口仍未弥合:「The end revenue from AI remains limited (on the order of tens of billions per year) relative to the scale of data center and energy investments (on the order of trillions over the coming five years)」;并记录 Dwarkesh 三联访谈后的「new consensus is that the AGI window will be in the 2030s, at earliest.」2026-07 经 TechCrunch 2026-07-09:2026 年 AI 基建支出 $1.5T → 需约 $3T 收入打平(⚠️ 该句原始载体待核,Sequoia 官网未见对应文章)。
Covello 的「$1tn 问题」。Goldman Sachs《Gen AI: too much spend, too little benefit?》2024-06-25(Wayback PDF 全文,主笔亲核):「We estimate that the AI infrastructure buildout will cost over $1tn in the next several years alone…So, the crucial question is: What $1tn problem will AI solve? Replacing low-wage jobs with tremendously costly technology is basically the polar opposite of the prior technology transitions…」(p.9)「AI technology is exceptionally expensive, and to justify those costs, the technology must be able to solve complex problems, which it isn’t designed to do.」实证细节:「AI can update historical data in our company models more quickly than doing so manually, but at six times the cost.」泡沫判断:「The NASDAQ declined around 70% between the highs of the dot-com boom and the founding of Uber.」可检验预言:「if important use cases don’t start to become more apparent in the next 12-18 months, investor enthusiasm may begin to fade.」(说于 2024-06)——到期检验:未兑现。18 个月后(2025 末)四大厂 capex 仍在加码、投资者热情未消退。Acemoglu 同报告的 TFP 账(「only 4.6% of all tasks will be impacted by AI…total factor productivity effects within the next decade should be no more than 0.66%…0.53%」)让位 AI 就业篇,此处只定位。
NANDA 的「95%」。MIT NANDA《The GenAI Divide》2025-07/08(报告 PDF 全文):执行摘要「Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.」方法论(p.2 逐字):「a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders」(注意:部分媒体写成 150/350,与原文不符);自述局限(p.6):「These figures are directionally accurate based on individual interviews rather than official company reporting.」漏斗:60% 评估→20% 试点→5% 投产;买 vs 建:「external partnerships…reached deployment ~67% of the time, compared to ~33% for internally built tools」;通用聊天机器人 pilot-to-implementation 约 83%——「95% 失败」针对的是定制化企业级系统,不是 ChatGPT 类工具。传播链:Fortune 2025-08-18 标题「95% of generative AI pilots at companies are failing」(报告原文无此句式,原文是「95% of organizations are getting zero return」)。方法论批评(Marketing AI Institute,The AI Show Ep.164 实录):「Zero is an extremely bold statement to make in any form of research」「This is not a viable, statistically valid thing.」——并记 MIT 后来把报告链接换成了申请表单。
官方层面警告。英格兰银行 FPC Record 2025-10-08(主笔亲核全文):「The risk of a sharp market correction has increased.」「equity market valuations appear stretched, particularly for technology companies focused on Artificial Intelligence (AI)」;「The market share of the top 5 members of the S&P 500, at close to 30%, was higher than at any point in the past 50 years.」CAPE 隐含盈利收益率「close to the lowest level in 25 years – comparable to the peak of the dot com bubble」;并点出 AI 特有下行通道:「Material bottlenecks to AI progress – from power, data, or commodity supply chains – as well as conceptual breakthroughs which change the anticipated AI infrastructure requirements…」Apollo 的 Torsten Sløk 2025-07-16:「The difference between the IT bubble in the 1990s and the AI bubble today is that the top 10 companies in the S&P 500 today are more overvalued than they were in the 1990s」。Michael Burry(2025-11):13F 显示 NVDA/PLTR 看跌敞口;X 帖(经 CNBC 通稿转引):「Understating depreciation by extending useful life of assets artificially boosts earnings – one of the more common frauds of the modern era.」估计五大厂 2026–2028 合计少提折旧约 $176B;Karp 反击(措辞待核)与 NVIDIA 分析师备忘录(原文未取得)双向标注。
DeepSeek 冲击的会计教训(2025-01-27,见第二节):「效率冲击→需求崩塌」的推断当日抹去英伟达 $589B 市值,但 Nadella 的 Jevons 推文与 SemiAnalysis 的判断(「Jevons is closer to reality, the models have already induced demand with tangible effects to H100 and H200 pricing」)被随后 18 个月的 capex 数字证实——效率提升没有减少算力需求,它改变了算力的用途结构(训练→推理)。
4. Stargate:承诺的解剖
OpenAI 官方公告,2025-01-21(Wayback 快照):「The Stargate Project is a new company which intends to invest $500 billion over the next four years building new AI infrastructure for OpenAI in the United States. We will begin deploying $100 billion immediately.」Musk 当日拆台(X,经新华社英文/AFP 多源一致转引):「They don’t actually have the money.」「SoftBank has well under $10B secured. I have that on good authority.」Altman 回击:「wrong, as you surely know… Want to come visit the first site already underway?」落地双向证据:加码侧——OpenAI 2025-07-22:Oracle 4.5GW、「over 5 gigawatts of Stargate AI data center capacity under development…over 2 million chips」「We now expect to exceed our initial commitment」;股权结构(二手分析):$500B 中承诺股权仅约 $52B,其余靠债务与供应商融资——Musk 质疑的结构性依据成立了一半。摩擦侧——The Information 2026-02-23:JV「stalled due to unresolved disputes among partners over site ownership and control」,未招聘员工、JV 名下未开发数据中心,OpenAI 转向双边交易;Bloomberg 2026-03-06 报道 Abilene 扩产取消,Oracle 在 X 否认(Tom’s Hardware 转引):「Recent media activity about the Abilene site are false and incorrect…」;Epoch AI 站点追踪(2026-04-17 更新):Abilene 运营容量估计约 0.6GW。
5. 第六层裁决
钱没有吓跑——「撞墙」论战后 18 个月,四大厂 capex 指引翻倍至约 $700B/年,NVIDIA 订单能见度从 $500B 上调到 $1T+(GTC 2026),Google token 用量两年 330 倍。账也没有合上——OpenAI 2025 年每挣 $1 花 $1.60、算力承诺自砍一半($1.4T→$600B)、Stargate JV 停摆、BoE/IMF/Sløk/Burry 的估值警告在案。两个检验分裂意味着:资本市场投票的是「scaling 叙事还有戏」,不是「scaling 已被证真」;而 OpenAI 的自我下修说明连最大的押注者也在给承诺表降杠杆。两个叙事各自的记账:讣告/泡沫派要记 Covello 的 12–18 个月预言到期未兑现、DeepSeek 恐慌被 Jevons 证伪;命定派要记 Cahn 的缺口公式从未被收入侧弥合、且 $3T 问题在 2026-07 仍悬在桌上(TechCrunch 同月文)。AI 泡沫之争是全库 EMH 篇的老朋友:价格真( capex 是真金白银)≠ 价值已证(收入缺口未闭)。
七、双向裁决层:讣告派四连错 × 命定派四连软
1. 讣告派四连错
其一,预言的停滞未发生。 2024-11 宣判「diminishing returns/plateaued」之后的 18 个月,是公开可核数字增长最快的 18 个月:SWE-bench 49%→93.9%、HLE 个位数→64.7%、ARC-AGI-2 4%→77.1%、METR 时间视界 1 小时→10 小时+。其二,「预训练已死」被宣判者本人部分收回。 Ilya 2025-11 在 X 上澄清 current scaling technology 会继续带来改进——他保留的只是「something important will still be missing」这一哲学性尾巴。其三,数据墙有缓冲。 「peak data」宣判后,业界 token 预算继续翻倍(15T→40T),Epoch 自己 2025 年判「unlikely to run out of data」;model collapse 之争以「累积可用」收场。其四,钱没跑。 如果讣告是定价信息,capex 曲线应该拐头;它翻倍了。
2. 命定派四连软
其一,「平滑可预测」失守于下游。 Kaplan 的承诺只覆盖 loss;Schaeffer 2024 证明下游指标可预测性系统性退化,而「撞墙」争论的恰恰是下游。其二,单轴放缓是事实。 Orion/GPT-4.5 被自家系统卡初版定性「not a frontier model」、30 倍定价一年下架 API 评估;行业换轴本身确认了单轴时代的终结——「换轴续命」与「从未有墙」不能同时为真。其三,承诺表反复滑移。 「there is no wall」(2024-11)之后是 GPT-5 发布混乱、4o 恢复、code red、承诺从 $1.4T 砍到 $600B;Altman 自己 2025-08 谈泡沫「smart people get overexcited about a kernel of truth」。其四,经济账未闭。 Cahn 的缺口公式从 $600B 滚到 $3T,收入侧从未弥合它;「more compute = better performance」在 RL 轴上已被 Meta 实测为 sigmoid——连新轴也不是幂律永动机。
3. 中间派的定调最接近落点
Nathan Lambert(2025-02-28,GPT-4.5 发布后 24 小时):「Scaling language models is not dead. Still…We’ve entered the era where trade-offs among different types of scaling are real.」「Now we really know what Ilya saw.」——讣告派的观察(单轴放缓)与命定派的行动(换轴继续)在同一个句子里共存。Nadella 的文献学注脚同样到位:「these are not physical laws. There are just empirical observations that hold true just like Moore’s Law did for a long period of time and so therefore it’s actually good to have some skepticism some debate.」黄仁勋 2026-03 的四轴清单(pretraining/post-training/test-time/agentic)则给这场争论提供了词源学落点:「scaling」这个词在 2023 年指一条轴,在 2026 年指四条轴——双方用旧词打新仗,是这场争论一半噪音的来源。
4. 母裁决
Scaling 没有撞墙;墙把这个词撞裂成了四条轴。 讣告派赢在观察(预训练单轴增益放缓,被官方文件确认)输在终审(全线停滞未发生、数据墙被缓冲、资本反手加码、宣判者本人部分收回);命定派赢在续命(推理时新轴真兑现,官方数字可核)输在外推(平滑可预测失守于下游、新轴已现 sigmoid 天花板、经济缺口未闭合、承诺表反复滑移)。对 2023 年定义的「预训练 scaling」,讣告判得七分;对 2026 年定义的「多轴 scaling」,讣告一分不值。这不是和稀泥——两边的对错落在不同的词义上,而词义本身在论战期间漂移了,这是本篇与「同名不同物」篇(2026-06-29)的共享病灶。
对称三向红线
- 不升格:不把新轴兑现读成「scaling 永动机」(ScaleRL 的 sigmoid、Overthinking 的边际递减、R1 的 limitation 都是一手自限);不把 capex 加码读成「市场永远对」(Cahn 缺口、BoE 警告、OpenAI 自我下修同案在卷);不把幂律读成物理定律(Amodei「misnomer」、Nadella「not physical laws」、Kaplan 自留「must flatten out eventually」退路——三位当事人各自拆过这个名号)。
- 不虚无化:讣告派的核心观察真——Orion 不及预期有官方文件确认、公共文本数据有限有定量估计、下游可预测性退化有论文证明、老基准饱和到 top2 只差 0.7%。Marcus 2022 年诊断的一部分(纯 scaling 不够)被时间部分背书:连最挺 scaling 的 Lambert 都写「Now we really know what Ilya saw」。讣告错在终审,不错在观察。
- 不污名:当事人修正立场是科学美德不是打脸素材——Ilya 的 walk-back、Epoch 的 v1→v2 自我修正、OpenAI 的 $1.4T→$600B 下修、METR 的自我勘误与「unreliable above 16 hrs」挂注,都应按内容记账而非按翻转记账;媒体报道与当事人原话的差异是传输层病(拼接引文、标题双版本、系统卡删行),注明版本即可,不上纲为造假;Marcus 的自我贴金「(and prescient)」与其论据的真实性分开记账。
关键来源
- 唯象地基:Kaplan 2020 arXiv:2001.08361|Chinchilla arXiv:2203.15556|Epoch 复现 arXiv:2404.10102|BNSL arXiv:2210.14891|Sorscher arXiv:2206.14486|Bahri arXiv:2102.06701|Schaeffer 2024 arXiv:2406.04391|Hägele arXiv:2405.18392
- 叙事考古:The Information 引爆稿|BNN Bloomberg 全文转载|Reuters 2024-11-11|The Verge Ilya 现场|Lex×Amodei 逐字稿|NVIDIA FY25Q3 电话会转录|Dwarkesh×Ilya 逐字稿|Lex×Huang 逐字稿|Marcus 宣判帖|Lambert GPT-4.5 定调
- 数据墙:Villalobos arXiv:2211.04325|Epoch 官方博客 2024|Epoch《Can AI Scaling Continue Through 2030?》|Epoch《AI in 2030》PDF|Shumailov Nature 2024|Gerstgrasser arXiv:2404.01413|Muennighoff arXiv:2305.16264|phi-4 arXiv:2412.08905
- 代际与基准:GPT-4.5 博文(Wayback)|GPT-5 博文|The Verge GPT-5 现场|Claude Opus 4.5|Project Glasswing|Gemini 3|Gemini 3.1 Pro|METR 时间视界|METR SWE-bench 水分研究|AI Index 2025
- 新轴:o1 博文|ARC Prize o3 分析|DeepSeek-R1 Nature|Snell arXiv:2408.03314|Meta ScaleRL arXiv:2510.13786|TTS 系统比较 arXiv:2512.02008|Overthinking arXiv:2604.10739|ARC-AGI-2|ARC Prize 2025 结果
- 经济账:Cahn《AI’s $600B Question》|Cahn《A Tale of Two AIs》|Goldman 报告(Wayback PDF)|MIT NANDA 报告 PDF|Stargate 公告(Wayback)|SemiAnalysis《DeepSeek Debates》|BoE FPC Record 2025-10|Apollo Sløk|TechCrunch 2026-07-09
灵魂句
Scaling 没有撞墙——墙把这个词撞裂成了四条轴。讣告派看见了裂缝就报了丧,命定派换了个轴就说裂缝不存在;而账单,两边都还没付。
机制裁决第 85 篇·对称双向第 80 篇·AI/认知谱系 scaling 侧·全库第 142 篇。调研纪律:六捆并行一手调研(标度律唯象地基/撞墙叙事考古/数据墙与合成数据/推理时新轴/模型代际与基准/经济资本与泡沫)+主笔亲核十四条承重引用(Kaplan 摘要与 §1 PDF/Chinchilla 摘要/Villalobos v2 摘要/Shumailov Nature 摘要(Crossref)/o1 博文与附录 A(Wayback)/o3/o4-mini 博文(Wayback)/Snell 摘要/ScaleRL 摘要/The Verge Ilya 现场/BNN Bloomberg 全文/Epoch 2030 全文/Cahn 全文/ARC Prize o3 页/Goldman PDF(pdftotext)/GPT-4.5·GPT-5·Gemini 3·Opus 4.5·Glasswing·METR 官方页),全部逐字一致;子代理纠错已并入正文(arXiv:2405.17818→2405.18392、arXiv:2406.15762→2408.00724、arXiv:2304.14177→2304.14108、TTT「8x」→6x、「largest and most knowledgeable」非博文原文、o1 系统卡无 AIME 跑分);拼接引文(Ilya 三合一)、标题双版本(WSJ)、系统卡删行(GPT-4.5)、成本三版本(ARC o3)均随文标注。二手材料(GPT-5.6 细节、Meta Watermelon、Grok 4.3、Anthropic 4.6/4.7 图表数字、Cahn $3T 载体、Karp 措辞)已逐处标注待核。提示注入:各捆均未发现可疑注入;OpenAI o1 博文抓取时曾遇正文提取器误吞 cipher 演示文本(已改用 Wayback),记为抓取事故非注入。不凭记忆,引用带链接;找不到的(NYT 原文、SemiAnalysis 部分付费稿、X 原推)如实交代获取路径。