目录
机制裁决第 120 篇 · 对称双向第 115 篇 · AI/认知谱系 · 全库第 178 篇 本篇是「合成数据与模型崩塌」的系统审计——问的是「模型吃自己的输出会不会死」。先读三句红线:
- 本篇不裁决任何公司、研究者或产品的对错,不给「该不该用合成数据」的工程建议,也不对 AI 未来发展做任何预测——只审「崩塌是定理还是叙事」。
- 「模型崩塌是数学定理」与「AI 即将死于自己的输出」是两件不同的事:前者有精确的成立条件,后者是把条件句去条件化后的末日叙事。本篇对称审计两个方向。
- 全篇承重句均给出可点击来源;正文英文逐字引用一律标注「一手逐字」(全文/页面取回)或「摘要逐字」(仅摘要取回,正文未取回),未取回原文的明确降档。
零、一句话裁决
模型崩塌是真定理,合成数据是真杠杆——真的不是「AI 即将死于自己的输出」和「合成数据是免费的无限燃料」这两层被声称的定论。 把我们的配比规则读成模型的宿命,是把一道工程选择题读成了一句死刑宣判。这篇报告要拆的,是这句话里每一个限定词的来处。
一、本篇测什么:三跳升格与五件可判定的事
本篇审的不是「AI 会不会崩溃」,而是「『AI 会崩溃』这句话是怎么从一条数学定理长成一则末日预言的」。拆成五件可判定的事:
- 定理真不真:在什么设定下崩塌是数学上不可避免的(Nature 2024 的 Gaussian 解析定理、三误差机制、跨设定实验)——这是守真锚。
- 条件在哪:「不可避免」的成立条件是什么,条件的边界被谁动过(arXiv 摘要 vs Nature 摘要的措辞差异、replace vs accumulate、1% 阈值)——这是条件账。
- 工程真相:业界真实用合成数据的方式与效果(phi-4 40% 配比与自消融、EMNLP 2025 千实验、R1 蒸馏)——这是工程账。
- 互联网事实:AI 内容在互联网与训练语料里的真实占比,以及「占比」这把尺本身可不可靠(检测器误报率)——这是互联网账。
- 叙事怎么长的:「模型吃自己」的每个隐喻词(dementia、curse、collapse、MAD、cannibalism)从哪里来,谁先用了末日修辞,作者自己后来怎么说——这是术语与末日账。
结构胎记:自噬跳 × 毒化跳 × 万能跳,附一条反向红跳。
- 自噬跳:把「AI 生成内容进入互联网与训练池」读成「模型在吃自己」——把数据分布的血缘关系读成个体的消化行为。单代模型不会「吃」任何东西;它是后代的训练数据里混入了前代输出的副本。
- 毒化跳:把「替换式递归训练必然崩塌」这一条件句读成「合成数据有毒」——条件被去条件化(与模拟篇析取去条件化同家族,但对象不同:那里是析取支被读成确定分支,这里是条件被读成无条件)。
- 万能跳(镜像):把「精心配比合成数据有效」读成「合成数据是免费的无限燃料、数据墙已解除」——同一条证据被两个方向各借一次。
- 反向红跳:尺子有噪声≠尺子无用(见第八节反虚无账)。
共用动作:把一句关于我们数据配比与采样规则的话,读成一句关于模型宿命的话。
与库内相关篇的分界:
- 与 scaling 篇(2026-07-23)分界写死:那篇把 model collapse 当作「数据墙」的一个论据写了四段(Shumailov/Gerstgrasser/Dohmatob×2/phi-4 的核心对峙与裁决小注「『合成数据必然毒化』与『合成数据随便用』都不立」已在那边);本篇把崩塌机制本身当对象——只称重不重做。那四篇的结论只在本篇第一至三节作为已核结论引用,增量全部在:机制数学的完整陈述、跨设定实验全谱、2025-2026 新实证、术语考古、互联网污染层(scaling 篇完全没碰)、工程全景与末日叙事层。并登记两处文献学更正:Gerstgrasser 篇非 NeurIPS 2024(dblp 仅 CoRR 收录,会议版为 ICML 2025《Collapse or Thrive?》,arXiv:2410.16713);phi-4 消融两句的实际位置是 §3.1「Data Composition in Pretraining」,不是 scaling 篇标注的「§4」。
- 与 AI 评测篇(2026-08-01)分界:那篇审「榜分=能力」的升格链;本篇审「模型吃自己的输出=模型死」的升格链——对象不同但共享同一母题:被读成宿命的那句话,本来是一句关于我们自己(测量/配比)的话。
- 与 AGI 篇(2026-08-01)分界:AGI 篇把 ARC-AGI 当完工判据;本篇不涉及「能力是否达到某水平」,只涉及训练数据回路。
- 与数字孪生篇(2026-07-26)分界:那篇审「模拟器=你」;本篇审「输出=食物」——共享「闭环」母题,但对象不同。
编号落位:机制裁决第 120 篇·对称双向第 115 篇·全库第 178 篇(2026-08-02 全库扫描确认未被占用)。
二、守真锚:崩塌是数学定理,且有跨设定实验复现
先清点定理本身。本库原则:审「声称」必须回到声称的原件。
2.1 Nature 2024:Gaussian 解析定理与「不可逆缺陷」
Shumailov et al. 2024, Nature 631, 755–759《AI models collapse when trained on recursively generated data》(arXiv:2305.17493,v1 提交 2023-05-27,Nature 版 Received 2023-10-20 / Accepted 2024-05-14 / Published 2024-07-24)[一手逐字,全文取回]。
arXiv v3 摘要(arXiv:2305.17493)关键句:
“We find that use of model-generated content in training causes irreversible defects in the resulting models, where tails of the original content distribution disappear. We refer to this effect as Model Collapse and show that it can occur in Variational Autoencoders, Gaussian Mixture Models and LLMs.” [一手逐字]
Nature 版摘要(s41586-024-07566-y)同位置关键句——注意新增一词:
“We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear. We refer to this effect as ‘model collapse’…” [一手逐字]
「indiscriminate」(不加区分地)一词只出现在 Nature 版摘要,arXiv v3 没有。这是全篇第一颗雷:定理的陈述在正式发表时被加上了限定词,而全球头条转述的正是这版带限定词的摘要,转述时又把限定词丢了。
论文对「不可避免」的完整主张(正文引言):
“We discover that indiscriminately learning from data produced by other models causes ‘model collapse’—a degenerative process whereby, over time, models forget the true underlying data distribution, even in the absence of a shift in the distribution over time. … Furthermore, we show that this process is inevitable, even for cases with almost ideal conditions for long-term learning, that is, no function estimation error.” [一手逐字]
注意:Nature 的主张是「即便没有函数估计误差(理想条件)也必然崩塌」——不是「无误差则不崩」。机制在三误差分类(正文「What is model collapse?」节)里说得很清楚:
“Statistical approximation error. This is the primary type of error, which arises owing to the number of samples being finite, and disappears as the number of samples tends to infinity.” [一手逐字]
“Functional expressivity error. … Even if we have perfect information about the data distribution (that is, infinite number of samples), model errors will be inevitable. However, in the absence of the other two types of error, this can only occur at the first generation.” [一手逐字]
“Functional approximation error. This is a secondary type of error, arising primarily from the limitations of learning procedures, for example, structural bias of stochastic gradient descent or choice of objective.” [一手逐字]
即:统计采样误差单独就足以致崩(样本有限→估计有偏→后代模型在更偏的估计上再估计);「Discrete distributions with exact approximation」一节更直接:模型崩塌仅由采样步骤的统计误差引起(”model collapse arises only because of statistical errors from the sampling step”)[一手逐字]。
核心解析定理(Theorem 3.1,Gaussian 设定,正文):
“Theorem 3.1 (Gaussian model collapse). Assume the original data are sampled from distribution D0 (not necessarily Gaussian), with non-zero sample variance. Assume Xn are fit recursively using the unbiased sample mean and variance estimators from the previous generation, Xjⁿ|μn, Σn ~ N(μn, Σn), with a fixed sample size. Then, E[W₂²(N(μn, Σn), D0)] → ∞; Σn → 0 as n → ∞, in which W₂ denotes the Wasserstein-2 distance between the true distribution and its approximation at generation n. In words, this implies that not only does the nth generation approximation diverge arbitrarily far from the original one but it also collapses to be zero variance as the number of generations increases, with probability 1.” [一手逐字]
一句话:在替换式递归、固定样本量、无偏估计的设定下,方差几乎必然收缩到零,近似分布与真实分布的 W₂ 距离几乎必然发散到无穷。这是定理,不是观点。
崩塌的正式定义(Definition 2.1):
“Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality. … We separate two special cases: early model collapse and late model collapse. In early model collapse, the model begins losing information about the tails of the distribution; in late model collapse, the model converges to a distribution that carries little resemblance to the original one, often with substantially reduced variance.” [一手逐字]
注意定义里的条件:「affecting generations of learned generative models」——崩塌是关于多代递归的命题,不是关于「训练数据里有一点 AI 内容」的命题。这是全篇结构性的分界线:互联网上单代模型混采 AI 内容,不满足「generations」条件。
论文还主动划清与近邻概念的边界(引言):
“We also briefly mention two close concepts to model collapse from the existing literature: catastrophic forgetting arising in the framework of task-free continual learning and data poisoning maliciously leading to unintended behaviour. Neither is able to explain the phenomenon of model collapse fully, as the setting is fundamentally different…” [一手逐字]
实验(Nature 版正文):OPT-125m 因果语言模型在 wikitext2 上微调,两种设定:
“Five epochs, no original training data. … We find that training with generated data allows us to adapt to the underlying task, losing some performance, from 20 to 28 perplexity points. Ten epochs, 10% of original training data preserved. … We find that preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance.” [一手逐字]
同一个实验里已经给出第一个条件句:保留 10% 原始数据→仅轻微退化。媒体报道里这个「10% 保留即可显著缓解」几乎从未被转述。
2.2 2022–2023:不是一个人发现的事
「模型吃自己的输出会退化」在 2022 年底到 2023 年有至少六组独立观察,时间线如下:
- Hataya, Bao & Arai(最早,arXiv:2211.08095,v1 提交 2022-11-15,ICCV 2023 pp. 20555–20565《Will Large-scale Generative Models Corrupt Future Datasets?》):
“Throughout experiments, we conclude that generated images negatively affect downstream performance, while the significance depends on tasks and the amount of generated images.” [一手逐字]
注意用词:contamination(污染),不是 collapse。
- Martínez et al. 2023(2023-03,合成数据迭代训练退化)。
- Shumailov et al. 2023(arXiv:2305.17493v1,2023-05-27)——原始标题是《Model Dementia: Generated Data Makes Models Forget》,摘要写「We call this effect model dementia」[一手逐字]。dementia(痴呆)是第一个医学隐喻。
- Alemohammad et al. 2023(arXiv:2307.01850,2023-07-04《Self-Consuming Generative Models Go MAD》)——第二个医学隐喻,且是疾病缩写:
“Repeating this process creates an autophagous (self-consuming) loop whose properties are poorly understood. … Our primary conclusion across all scenarios is that without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease. We term this condition Model Autophagy Disorder (MAD), making analogy to mad cow disease.” [一手逐字]
MAD 类比的是疯牛病——「吃自己的同类得病」。三种循环定义(§1.2):
“The fully synthetic loop, wherein the training dataset for each generation’s model consists solely of synthetic data sampled from previous generations’ models.” [一手逐字] “The synthetic augmentation loop, wherein the training dataset for each generation’s model is trained on a dataset Dt = (Dr, Dst) consisting of a fixed set of real data Dr sampled from Pr plus synthetic data Dst from models from previous generations.” [一手逐字] “The fresh data loop, in which each model Gt for t ≥ 2 is trained on a dataset Dt = (Drt, Dst) consisting of a fresh set of real data Drt drawn independently from Pr plus synthetic data Dst from models from previous generations.” [一手逐字]
(fresh data loop 即后来 Gerstgrasser 版「accumulate」的前身;实验模型含 StyleGAN2、DDPM、WGAN 等。)
- Bohacek & Farid(arXiv:2311.12202,2023-11-20,Stable Diffusion v2.1):
“We show that when retrained on even small amounts of their own creation, these generative-AI models produce highly distorted images. We also show that this distortion extends beyond the text prompts used in retraining, and that once affected, the models struggle to fully heal even after retraining on only real images.” [一手逐字]
- Bertrand et al.(arXiv:2310.00429,2023-09-30,ICLR 2024《On the Stability of Iterative Retraining of Generative Models on their own Data》)与 Briesch, Sobania & Rothlauf(arXiv:2311.16822,2023-11-28,从图像域把「self-consuming loop 降低质量与多样性」的结论带到 LLM 域)[一手逐字]。
这一节的小结:崩塌在多代递归设定下是真定理、被多团队独立发现、跨 GMM/VAE/GAN/扩散/LLM 复现;但它的每个版本都自带条件句(recursive / replace / self-consuming / without enough fresh real data)——条件句是定理的一部分,不是可以被摘掉的外包装。
三、机制账:为什么会崩
3.1 三个误差的放大回路
把 2.1 的三误差串成回路:第一代模型在有限样本上估计分布(统计误差 ε1)→ 表达能力有限(表达误差 ε2)→ 学习过程本身有偏(逼近误差 ε3)→ 第二代用含误差的合成数据再估计(ε1 在更偏的分布上再来一次)→ 误差逐代复合。Nature 的关键洞见是统计误差单独就够——即便表达与逼近误差都为零(「exact approximation」),有限样本的估计方差也会逐代累积直至方差收缩为零。
1D Gaussian 的递推结构(arXiv v3 正文)给出直觉:
“From Equation (3), we see that even after the first approximation, the distribution of Xji is no longer normal, it follows a variance-gamma distribution.” [一手逐字]
即:第一代之后分布就不再是正态(变分伽马分布),尾部开始变形——「tails of the original content distribution disappear」的机制起点。
3.2 为什么是「尾巴」先消失
Definition 2.1 里的 early collapse 说的是「begin losing information about the tails of the distribution」。机制上:有限样本估计对低概率区域(长尾)的估计最差(样本少→方差大),而合成数据又从这些已经失真的估计里采样,尾部逐代被截断。这解释了为什么崩塌首先表现为多样性下降(recall 丢失)而非平均质量的崩塌——Alemohammad 的 MAD 结论把 precision/recall 分开报(「quality (precision) or diversity (recall) progressively decrease」[一手逐字],见 2.2)。
3.3 与 GAN mode collapse 的血缘
「model collapse」这个命名直接继承 GAN 的 mode collapse。Shumailov v3 脚注 1 自述:
“The name is inspired by the Generative Adversarial Networks (GAN) literature on mode collapse, where GANs start producing a limited set of outputs that all trick the discriminator. Model Collapse is a process whereby models eventually converge to a state similar to that of a GAN Mode Collapse.” [一手逐字]
同脚注还有改名经过的官方自述(这是术语考古的承重原件):
“The original version of this paper referred to this effect as ‘model dementia’, but we decided to change this following feedback that it trivialised the medical notion of ‘dementia’ and could cause offence.” [一手逐字]
两颗雷同时落袋:① model collapse 的名字来自 GAN 文献的近似类比(「终态相似」);② v1 的医学隐喻(dementia)是作者自己决定换掉的——因为 trivialise 医学概念。这为第六节的末日叙事账埋下伏笔:作者团队对隐喻有自觉,但隐喻的扩散不受作者控制。
四、条件账:什么设定下崩,什么设定下不崩
4.1 replace vs accumulate:全篇最重要的限定词
Gerstgrasser et al. 2024《Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data》(arXiv:2404.01413)摘要逐字:
“Recent investigations into model-data feedback loops proposed that such loops would lead to a phenomenon termed model collapse, under which performance progressively degrades with each model-data feedback iteration until fitted models become useless. However, those studies largely assumed that new data replace old data over time, where an arguably more realistic assumption is that data accumulate over time. … We confirm that replacing the original real data by each generation’s synthetic data does indeed tend towards model collapse, then demonstrate that accumulating the successive generations of synthetic data alongside the original real data avoids model collapse; these results hold across a range of model sizes, architectures, and hyperparameters.” [一手逐字]
理论结果(Theorem 2,线性模型、累积设定):
“For an n-fold synthetic data generation process with T ≥ d + 2 samples per iteration and isotropic features (Σ =def Id), the test error for the ridgeless linear predictor ŵn learned on the accumulated data up to iteration n is given by: E_test^Accum(ŵn) = σ²d/(T−d−1) · (Σ_{i=1}^n 1/i²) ≤ σ²d/(T−d−1) × π²/6” [一手逐字]
π²/6 = ζ(2)(巴塞尔级数)——累积设定下测试误差有与代数无关的有限上界,而替换设定下线性增长(E_test^Replace ∝ n)。作者在讨论里把话说满又收回:
“Together, these results strongly suggest that the ‘curse of recursion’ may not be as dire as had been portrayed – provided we accumulate synthetic data alongside real data, rather than replacing real data by synthetic data only.” [一手逐字]
保留条款(作者自报,两处):VAE 实验中累积误差仍随代数增长(”albeit much more slowly”);且:
“Lastly, it is worth noting that ‘model collapse’ – as a term of art – has been used in various ways by various researchers; so care is required in comparing claims across articles. In reviewing the literature, we identified at least four related phenomena: (0) unbounded test error blowup (as here); (1) modal collapse — collapse to one (or a few) modes; (2) collapse to uniformity; and (3) amplification of artifacts introduced by models fit to previous synthetic data.” [一手逐字]
作者自己警告:『model collapse』是行话(term of art),至少有四义,跨文比较必须小心——这句话本身就该被全文引用。
会议版收录状态更正:dblp 收录 Gerstgrasser 2024 仅 CoRR(arXiv),无 NeurIPS 2024 收录记录;该课题的会议版是 ICML 2025《Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World》(arXiv:2410.16713,PMLR 267:29469–29494),其摘要把「崩塌与否」直接拆成三种训练工作流 × 三个任务设定的条件句:
“we report experiments on three ways of using data (training-workflows), across three generative model task-settings (multivariate Gaussian estimation, kernel density estimation, and language-model fine-tuning) to further confirm the possibility of containment: (a) we confirm that the training-workflow of replacing all real data by successive generations of purely synthetic data indeed suffers model collapse in all task-settings studied; (b) we consider the training-workflow of accumulating synthetic data alongside real data … models remain stable and their test losses do not diverge under this training-workflow; (c) we consider a training-workflow where real and synthetic data accumulate together but successive generations of pretraining are constrained to use fixed-size data subsets each generation. In this workflow, we observe slow and gradual rather than explosive degradation of test loss performance across generations.” [一手逐字]
4.2 Dohmatob 两篇:1% 阈值与它的设定边界
Dohmatob et al.《Model Collapse Demystified: The Case of Regression》(arXiv:2402.07712)在高维回归设定给出解析刻画,关键结论句:
“(b) In contrast, for the case where the Xn’s are independent, the increase in bias term grows with n, leading to ‘catastrophic’ model collapse (Theorem 4.9).” [一手逐字]
即:崩塌与否还取决于噪声结构——相依输入时偏置项不随代数增长,独立输入时才灾难性崩塌;且每代数据量固定是线性退化的原因(Remark 4.2:若数据量随代数增长则显著缓解)。
《Strong Model Collapse》(arXiv:2410.04840)摘要:
“Our results show that even the smallest fraction of synthetic data (e.g., as little as 1% of the total training dataset) can still lead to model collapse: larger and larger training sets do not enhance performance.” [一手逐字]
设定边界(必须随文带上):supervised regression 设定(BabiStories×GPT-2-small 124M + MNIST 回归损失);「1%」指单次混合的总训练集占比;结论是「只要合成占比不趋于零(p₂ 不→0),缩放训练集无法挽回」:
“Result #1: Strong Model Collapse. First, we establish a robust negative result which shows that model collapse generally persists even when mixing real and synthetic data, as long as the fraction of training data which is synthetic does not vanish.” [一手逐字]
这条是「毒化跳」最硬的材料:1% 就够——但它带三个限定词(回归设定、混合比例不趋于零、单次混合)。媒体转述时三个限定词通常一个都不带。
〔2026-08-03 增量·版本差〕「1%」还有一层版本学:arXiv 摘要作「as little as 1% of the total training dataset」(且含原文拼写 existance),而 ICLR 2025 Spotlight 会议版摘要收紧为「as little as 1 per 1000」、措辞改为「establish a strong form」(OpenReview API 逐字,增量方亲核)——会议版把阈值又压低了一个数量级,引用时须注明版本。
4.3 2025–2026:条件账继续被改写
- 理论侧反驳:Barzilai & Shamir《When Models Don’t Collapse: On the Consistency of Iterative MLE》(arXiv:2505.19046,NeurIPS 2025):
> “we establish non-asymptotic bounds showing that collapse can be avoided even as the fraction of real data vanishes. … this result highlights that model collapse is not inevitable, even when T → ∞ and the fraction of real data vanishes.” [一手逐字] - 立场文(方法论批判):Schaeffer et al.《Position: Model Collapse Does Not Mean What You Think》(arXiv:2503.03150,venue 存疑见诚实空位 11):
> “Industry leaders, premier research journals and popular science publications alike have prophesied catastrophic societal consequences stemming from model collapse. In this position piece, we contend this widespread narrative fundamentally misunderstands the scientific evidence. We highlight that research on model collapse actually encompasses eight distinct and at times conflicting definitions of model collapse, and argue that inconsistent terminology within and between papers has hindered building a comprehensive understanding of model collapse.” [一手逐字] - 实证侧(LLM 多代):《Knowledge Collapse in LLMs》(arXiv:2509.04796,2025-09-05):
> “we define knowledge collapse as a distinct three-stage phenomenon where factual accuracy deteriorates while surface fluency persists, creating ‘confidently wrong’ outputs… we demonstrate that collapse trajectory and timing depend critically on instruction format.” [一手逐字] - 实证侧(选择性反馈逆转):《The Anti-Ouroboros Effect》(arXiv:2509.10509,2025-09-02,5 代递归微调):
> “Across five generations, a quality-filtered condition improved by 6.6% in ROUGE-L F1 score, whereas an unfiltered control degraded by 3.5% and a random-filter control degraded by 4.2%.” [一手逐字] - 检测/重采样缓解:《Machine-generated text detection prevents language model collapse》(EMNLP 2025):
> “We demonstrate that it not only prevents model collapse but also improves performance compared to training on purely human data, underscoring the benefit of synthetic samples and the importance of data curation.” [一手逐字] - 验证器缓解:《Escaping Model Collapse via Synthetic Data Verification》(arXiv:2510.16657):外部可靠验证器存在时,合成重训不崩塌(摘要级)[一手逐字]。
- 观测侧:ChatGPT 各版本输出多样性的纵向观测(arXiv:2603.12683,2026-03):
> “Our findings indicate a measurable decline of recent ChatGPT releases’ ability to produce varied text, even when explicitly prompted to do so, by setting the temperature parameter to one.” [一手逐字]
(注意:这是「生成文本多样性下降」的观测,不是「模型在递归训练中崩塌」——两者的关系需要另外的因果链,本篇不建立。)
4.3.1 增量补遗(2026-08-03):理论三件与新实验四件
〔本节为 2026-08-03 增量,材料来自增量方理论地基捆与批判捆+主笔亲核(arXiv 摘要/OpenReview venue 逐字)〕
理论三件:
- Seddik et al.《How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse》(arXiv:2404.05090)——纯采样误差崩塌的最干净模型(词表上的类别分布直方图 LM,无函数逼近误差):
“Specifically, we demonstrate that model collapse cannot be avoided when training solely on synthetic data. However, when mixing both real and synthetic data, we provide an estimate of a maximal amount of synthetic data below which model collapse can eventually be avoided.” [摘要逐字]
Theorem 1 给出全坍缩期望代数 E[T] 介于 O(n) 与 O(n²) 之间(n 为每代样本数);Proposition 1:坍缩方向以初始概率 pᵢ 落在 token i 的 Dirac 分布——理论上最干净地坐实了 Shumailov「统计近似误差是主因」的机制定位;但模型是 0 阶直方图 LM(无序列结构),不能外推到困惑度或下游任务。
- Ferbach et al.《Self-consuming Generative Models with Curated Data Provably Optimize Human Preferences》(arXiv:2407.09499,NeurIPS 2024):
“We prove that, if the data is curated according to a reward model, then the expected reward of the iterative retraining procedure is maximized.” [摘要逐字]
同一篇的 Theorem 2.1 证明学习分布「will lose diversity and collapse to the highest reward samples」——「策展即偏好优化」与「多样性坍缩到最高奖励层」是同一枚硬币的两面;且实验显示该过程「amplifies biases of the reward model」(CIFAR-10 以分类器置信度策展 → 偏向 airplane 类)。引用「provably optimize human preferences」时必须带坍缩面,否则就是只摘半句。
- Dohmatob et al.《A Tale of Tails: Model Collapse as a Change of Scaling Laws》(arXiv:2402.07043,ICML 2024)——崩塌在此被操作化为 scaling law 形态改变(平台期/随代数平移/grokking),而非困惑度爆炸:
“Thus, as soon as T ≳ kᵝ, the AI-generated sample size T ceases to be a ‘scalable’ resource: collecting more AI-generated samples will not improve the performance of the downstream model, i.e performance plateaus and we lose scaling.”(Theorem 2.1)[一手逐字]
且 Theorem 3.2 给出反方向:任意小比例 π>0 的真实数据混入,测试误差最终恢复随 T 的 scaling(grokking)——「真实数据占比」而非「绝对量」是分水岭。
新实验四件(对 4.3 节的补充):
- 剂量-漂移曲线:Kovač et al.《Recursive Training Loops in LLMs》(arXiv:2504.03814,EMNLP 2025 Oral)——1–2B 模型 × 5 数据集 × 20 代 × 5 seeds、1,600 条链的迄今最大受控递归实验:
“Chains with r = 1/16 exhibit almost no shifts and chains with ratios r = 1/8 and r = 1/4 exhibit increasingly more shift. This seems to plateau at r = 1/2” [一手逐字] “Lexical diversity is found to amplify these shifts, while semantic diversity and data quality mitigate them” [一手逐字]
——漂移随合成占比单调递增但在 r=1/2 处平台化;跨域影响高度模块化(21 个显著预测子中仅 3 个跨域)。
- 验证器的边界(对「筛选即防坍缩」的首次严格反例):Qiao et al.《When Sample Selection Bias Precipitates Model Collapse》(arXiv:2606.13732,2026-06):
“in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection itself becomes biased. … turning from a safeguard against collapse into a mechanism that precipitates it. We theoretically prove that such siloed selection accelerates collapse and induces power-law diversity decay.” [摘要逐字]
——筛选防坍缩的前提是验证器参照分布足够全局;数据孤岛(医疗联盟、金融机构)场景下筛选反而剪尾。
- 后训练特质维度:Roe et al.《Iterative Finetuning is Mostly Idempotent》(arXiv:2605.01130,2026-05):
“In the SFT and SDF settings, traits mostly decay or remain constant so that further finetuning cycles do nothing. In rare cases when amplification occurs, it generally comes at the cost of coherence.” [摘要逐字]
——特质(persona/bias)维度上自训循环多为衰减而非放大;与「多样性坍缩」不同物,是反方向证据。
- 代码域:《When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs》(arXiv:2606.28438,2026)——模型自评门控被证明会退化:
“AI self-gating degenerates to ungated self-training under a self-confirming acceptance condition” … “stable recursive code LLM training requires exogenous verification rather than model-coupled self-review.” [一手逐字]
——与 4.3 的验证器缓解形成精确互补:能防坍缩的是外生验证器,不是模型自评。
4.4 条件账小结(设定边界清单)
| 条件 | 崩塌成立? | 出处 |
|---|---|---|
| 替换式(replace)递归 | 成立,误差线性增长 | 2404.01413 式(4);2402.07712 Thm 4.9 |
| 累积式(accumulate) | 不成立,π²/6 有限上界 | 2404.01413 Thm 2 |
| 累积但每代固定大小子集 | 缓慢退化、不爆炸 | 2410.16713(ICML 2025) |
| 无函数估计误差(理想条件) | 仍崩(仅统计误差即可) | Nature 正文 |
| 独立输入噪声 vs 相依 | 独立→灾难性;相依→偏置不增长 | 2402.07712 §4.6/4.9 |
| 数据量每代固定 vs 增长 | 固定→线性退化;增长→显著缓解 | 2402.07712 Remark 4.2 |
| 合成占比趋于零 vs 不趋于零 | 不趋于零→强崩塌(1% 即够) | 2410.04840 Result #1 |
| VAE(图像)累积 | 误差仍增长(慢得多) | 2404.01413 §2.3 |
| 采样温度 | 0.3 比 1.0 退化更快 | 2404.01413 §2.1 |
| 迭代 MLE 一致性 | 可避免(真实数据占比→0 也稳) | 2505.19046 |
| 外部验证器 | 不崩塌(收敛到验证器中心) | 2510.16657 |
| 质量过滤(selective feedback) | 5 代反而 +6.6%(对照 −3.5%) | 2509.10509 |
这一节的小结:把「崩塌不可避免」这六个字换成它真正的主语——「替换式递归 + 固定样本 + 无验证」设定下崩塌不可避免。限定词不是文章的细节,是文章的主体;媒体转述丢了限定词,定理就变成了预言。
五、工程账:合成数据是怎么被真正使用的
5.1 蒸馏谱系:合成数据是成熟工程杠杆
合成数据不是 2024 年崩塌论文之后才有的东西。从 2022 年底起,它是主流训练管线的标准零件:
- Self-Instruct(arXiv:2212.10560,2022-12):
> “We introduce Self-Instruct, a framework for improving the instruction-following capabilities of pretrained language models by bootstrapping off their own generations. Our pipeline generates instructions, input, and output samples from a language model, then filters invalid or similar ones before using them to finetune the original model.” [一手逐字]
——自举(bootstrapping)+ 过滤(filters invalid or similar ones):从第一天起,工程做法的关键就是质量控制,不是「喂多少合成数据」。 - WizardLM(arXiv:2304.12244,2023-04):Evol-Instruct 用 LLM 把指令逐步改复杂。
- Orca(arXiv:2306.02707,2023-06):学 GPT-4 的推理痕迹(explanation traces);Orca 2(arXiv:2311.11045)的转向句是早期自觉:
> “We contend that excessive emphasis on imitation may restrict the potential of smaller models.” [一手逐字] - phi 系列(微软):phi-1(arXiv:2306.11644,6B 精选网页 token + GPT-3.5 合成的 1B 教科书式 token);phi-1.5(arXiv:2309.05463);phi-2(模型卡,250B token 混合 GPT-3.5 合成 + 过滤网页)。「教科书质量」的合成数据是 phi 系列的设计核心。
- phi-4(arXiv:2412.08905,2024-12)——全篇最重要的工程承重墙:
- 摘要:
> “While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go beyond distillation.” [一手逐字] - 正文(引言):
> “Synthetic data constitutes the bulk of the training data for phi-4 and is generated using a diverse array of techniques, including multi-agent prompting, self-revision workflows, and instruction reversal.” [一手逐字] - 配比(§3.1 Table 5):Web 15% / Web rewrites 15% / Synthetic 40% / Code 20% / Acquired 10% [一手逐字]
- 消融保留条款(§3.1,双向审查关键材料,两句同段):
> “We note two key observations. • Web datasets showed small benefits on reasoning heavy benchmarks. Prioritizing more epochs over our synthetic data led to better performance with respect to adding fresh web tokens. • Models trained only with synthetic data underperformed on the knowledge-heavy benchmarks and demonstrated increased hallucinations.” [一手逐字]
——同一段两句话:一句支持「多轮次合成优于加新网页 token」,一句自我设限「纯合成在知识密集榜掉分+幻觉增加」。这就是「杠杆」的形状:合成数据有用,但「只用合成数据」和「新鲜网页优先」两头都错,工程答案是配比(40%)。 - Cosmopedia(HF,Mixtral-8x7B 生成,30M 文件/25B token):官方载体是 HF 博客(huggingface.co/blog/cosmopedia)而非 arXiv:
> “we introduce Cosmopedia, a dataset of synthetic textbooks, blog posts, stories, posts, and WikiHow articles generated by Mixtral-8x7B-Instruct-v0.1. It contains over 30 million files and 25 billion tokens, making it the largest open synthetic dataset to date.” [一手逐字] - Llama 4(Meta,2025-04):官方博客确认 Behemoth 蒸馏出 Maverick(codistillation),预训练 30T+ token,但未披露合成数据占比(模型卡仅确认微调/安全环节用了合成数据)[一手逐字,经官方博客与模型卡]。
5.2 自进化谱系:RL 自训练是「另一种合成数据」
- Self-Rewarding LMs(arXiv:2401.10020,2024-01):LLM-as-a-Judge 自打分迭代 DPO,自举(iteration 中自己产生的数据继续训练自己)。
- DeepSeek-R1 / R1-Zero(arXiv:2501.12948,2025-01)——自进化最硬的一手:
> “Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories.” [一手逐字]
R1-Zero 的 AIME 轨迹:从 15.6% 到 77.9%(正文)[一手逐字]——纯 RL 自对弈式训练在推理任务上大幅提升(规则可验证奖励的 AlphaZero 式自训练,不是崩塌式递归)。
蒸馏策略(正文):
> “Specifically, we fine-tune open-source foundation models such as Qwen … and LLaMA … using a curated dataset comprising 800,000 samples generated with DeepSeek-R1.” [一手逐字]
> “We find that models distilled from high-quality teacher outputs consistently outperform those trained directly on human-generated data, corroborating prior findings on the efficacy of distillation.” [一手逐字] - Altman 的官方口径(CNBC Squawk Box 采访转录,2025-08-08,页面自标 unofficial transcript):
> “we really started using synthetic data. So, the previous generation of the model is teaching the next generation of the model. And as these models get smarter, the data that they can create can be, you know, really quite interesting and helpful.” [一手逐字,转录稿]
对照事实:GPT-5 系统卡(官方 PDF,2025-08-07)全文synthetic出现 0 次 [一手逐字]——CEO 采访与官方文档之间的披露差,本身是「合成数据占比」没有官方口径的证据。 - Anthropic Claude Sonnet 5 系统卡(官方 PDF):训练数据含「synthetic data generated by other models」[一手逐字]——确认合成数据在训练混合里,配比未披露。
- 扎克伯格(Dwarkesh Patel 采访,2024-04-18,官方文字稿):
> “I do think in the future it seems quite possible that more of what we call training for these big models is actually more along the lines of inference generating synthetic data to then go feed into the model. I don’t know what that ratio is going to be but I consider the generation of synthetic data to be more inference than training today.” [一手逐字] - Ilya Sutskever(NeurIPS 2024 演讲):「数据是 AI 的化石燃料」「peak data」——但注意:无官方逐字稿,只有两份社区转录(内容一致),本篇按转述处理,不承重为逐字引用(详见诚实空位)。
5.3 配比科学:2025–2026 的大规模实证
合成数据的配比问题从「要不要用」变成「用多少」:
- EMNLP 2025《Demystifying Synthetic Data in LLM Pre-training》(arXiv:2510.01631,>1000 个 LLM 训练实验、>100k GPU 小时):
> “we found pre-training on rephrased synthetic data alone is not faster than pre-training on natural web texts; while pre-training on 1/3 rephrased synthetic data mixed with 2/3 natural web texts can speed up 5-10x (to reach the same validation loss) at larger data budgets.” [一手逐字]
> “‘Good’ ratios of synthetic data in training data mixtures depend on the model size and data budget, empirically converging to ~30% for rephrased synthetic data.” [一手逐字]
> “training on rephrased synthetic data shows no degradation in performance in foreseeable scales whereas training on mixtures of textbook-style pure-generated synthetic data shows patterns predicted by ‘model collapse’.” [一手逐字]
——三条合起来是工程账的地基:纯合成不更快;1/3 配比快 5–10 倍;改写式合成无退化而「教科书式纯生成」出现崩塌模式。合成数据的类型和配比共同决定方向。 - 《Synthetic Data Proportions》(arXiv:2510.05133,2025):
> “models maintain stable performance with up to 20% synthetic data, but degradation accelerates beyond 30%; larger models (6.9B-12B) show greater robustness to synthetic data than smaller models (410M-1.4B); calibration degradation precedes accuracy loss” [一手逐字] - Kazdan et al. ICML 2025《Collapse or Thrive?》(PMLR 267):
> “when real data are scarce, there exists an optimal amount of synthetic data that are helpful” [一手逐字]
——「真实数据越稀缺、合成数据越有用」的机制性表述,且有最优量。
5.4 纯合成的上限实证(反方材料,不虚无化)
- 知识密集:Knowledge Collapse(见 4.3)——事实精度恶化而表面流利度保持(「confidently wrong」)。
- 多语言:ACL 2026(2026.acl-long.1002):合成数据注入世界知识有效,但直接训练「degrades native semantic fluency」[一手逐字];低资源语言差距更大(arXiv:2506.12158:威尔士语最优生成设定仍比金标差 −11.53%)[一手逐字]。
- 预训练下游域:EMNLP 2025(同 5.3):教科书式纯合成在多个下游域损失更高(尤其小数据预算)[一手逐字]。
- 无控环境:Barbaro et al. 2025(HAL):FineWeb dump 中约 16% 文本被检测为 GenAI;去除合成数据后「saved more than 40% in computation time, while obtaining superior results on average」[一手逐字]——无控环境 vs 受控配比的对照。
- 成功案例(平衡):MetaMath(arXiv:2309.12284)GSM8K 66.4%/MATH 19.4%,MetaMath-70B 82.3% 略超 GPT-3.5-Turbo;WizardCoder(arXiv:2306.08568)在 HumanEval 超 Claude/Bard;DeepSeek-Prover(arXiv:2405.14333)miniF2F 46.3%(64 样本)对 GPT-4 的 23.0%——数学/代码/形式证明这三类「答案可验证」的任务上合成数据明确有效,而这恰是「验证器存在时不崩」(4.3 节 2510.16657)的工程侧印证。
这一节的小结:工程世界里「合成数据有毒」和「合成数据是免费燃料」都没有对应物。真实形状是一个旋钮:类型(改写 vs 教科书式)、配比(约 20–40%)、规模(大模型更韧)、验证(可验证任务收益大)四个维度决定方向。崩塌论文给了旋钮的数学,工程界一直在转旋钮。
5.5 增量补遗(2026-08-03):可验证域、对齐域、开源指令域与非 LLM 域
〔本节为 2026-08-03 增量,材料来自增量方工程实证捆+主笔亲核〕
- 可验证域的范式案例:AlphaGeometry(Trinh et al., Nature 625, 476–482,DOI):
> “By using existing symbolic engines on a diverse set of random theorem premises, we extracted 100 million synthetic theorems and their proofs, many with more than 200 proof steps, four times longer than the average proof length of olympiad theorems.” [一手逐字,正文 Main 节]
摘要自称「sidesteps the need for human demonstrations」;IMO-AG-30 测试集解出 25/30(前 SOTA 吴方法 10 题;仅用 20% 训练数据仍解 21 题)。合成样本的正确性由符号演绎引擎保证,语言模型只学最难的辅助构造——「验证器存在时合成数据无上限」的最强工程证据。作者自留限定:「AlphaGeometry operates with a much lower-level toolkit for proving than humans do, limiting the coverage of the synthetic data, test-time performance and proof readability」[一手逐字]。 - 对齐域的极限配比:NVIDIA Nemotron-4 340B(arXiv:2406.11704):
> “Notably, over 98% of data used in our model alignment process is synthetically generated, showcasing the effectiveness of these models in generating synthetic data.” [摘要逐字]
——分母是对齐过程(SFT+偏好训练)而非预训练;NVIDIA 明确把这组模型定位为给社区生成合成数据用的教师模型。 - 开源指令域:Magpie(arXiv:2406.08464)从 Llama-3-Instruct 自抽取 400 万指令-响应对、精选 30 万:
> “models fine-tuned with Magpie perform comparably to the official Llama-3-8B-Instruct, despite the latter being enhanced with 10 million data points through supervised fine-tuning (SFT) and subsequent feedback learning.” [摘要逐字]
——30 万对 1,000 万持平。SmolLM2(arXiv:2502.02737)的分工表述:预训练 11T token 中 Cosmopedia v2 合成教科书占 4%(退火期);指令侧用 Llama-3.1-405B 生成 Magpie-Ultra 1M 三轮对话,再以 8B 安全/质量模型与 ArmoRM 过滤——「AI 标注用于筛选、AI 生成用于补强」 [一手逐字]。 - 非 LLM 域的打折证据:RoentGen(Chambon et al., arXiv:2211.12737,胸片 latent diffusion):
> “Fine-tuning this model on a fixed training set and using it as a data augmentation method, we measure a 5% improvement of a classifier trained jointly on synthetic and real images, and a 3% improvement when trained on a larger but purely synthetic training set.” [摘要逐字]
——混合 +5%、纯合成 +3%:与 LLM 侧「合成做增量、纯合成打折」同型;文本编码器还顺带获得领域知识(气胸表征 +25%)。
六、互联网账:AI 内容真的在占领互联网吗
「互联网正在被 AI 内容淹没」是自噬跳的事实层。事实层要精确到口径——因为同一时点不同口径的数字可以差一个数量级。
6.1 占比的七把尺(按口径分列)
| 口径 | 数字 | 时点 | 出处 |
|---|---|---|---|
| 新网页「含 AI 成分」 | 74.2%(纯 AI 2.5%、混合 71.7%) | 2025-04 | Ahrefs(90 万新页面样本,bot_or_not 检测) |
| 新文章「主要 AI 生成」 | 2025Q1 49.6% vs 50.4%;2025Q4 50.9% 超越;2026Q1 49.9% | 2020-01 至 2026-03 | Graphite(Common Crawl 随机抽样 55.4k 文章,三检测器平均) |
| 活动网页文本来源 | ≥30%,可能近 40% | 2025-03 | Spennemann(关键词频次法,作者自述 40% 可能偏高) |
| 新上传网站 | 约 35% AI 生成/辅助 | 2025 上半年 | Dolezal et al. 2026(Wayback 分层抽样,Pangram v3) |
| 训练语料站点级(LLM-dominant) | 2.1% → 29.4% | 2022H2 → 2025H1 | DeGenTWeb(Common Crawl 94,908 站分类) |
| 网页「含 AI」早期估计 | 0.0185% → 1.57%(8,362% 增长) | 2022-12 → 2024-03 | Copyleaks 2024(100 万页面样本;官方发布页已清档,两源交叉核对) |
| Amazon 新电子书检出 AI | 2023 月峰 30% → 2024 45% → 2025 超 60% | 2025 年底 | NBER WP 34777(5 万+随机样本,Pangram) |
(说明:Copyleaks 官方发布页已下架,数字经 TechNewsWorld 报道与其官方博客两源交叉核对,降为 ◐ 档;其余均直取一手。)
纵横向参考:维基百科新建英文条目检出显著 AI 内容 4.36%(2024-08,arXiv:2410.08044,1% FPR 阈值校准,作者明示为下限);arXiv/bioRxiv/Nature 论文摘要与引言中被 LLM 显著修改的句子占比,CS 至 17.5%(arXiv:2404.01268);NewsGuard 追踪的 AI 内容农场从 2023-05 的 49 个 → 2023-06 150 个 → 2024-08 近 1,000 个 → 2024-11 1,121 个 → 2026-06 3,749 个(NewsGuard AI Tracking Center,官方页直取 2026-06 时点数字)。
关键观察:「AI 内容在增长」是七个口径共同指向的事实(从 1.57% 到 74.2% 都有,取决于口径);「AI 内容占比」本身不是「模型在吃自己」的证据——它只是说明互联网这个池子里新注入的 AI 文本变多了。真正决定模型命运的是训练管线怎么从这个池子采样(递归?替换?配比?验证?)——即第四节的设定条件。
6.2 尺子的反身性:检测器自己靠不住
占比数字全部依赖 AI 检测器。检测器的可靠性:
- Liang et al. 2023(arXiv:2304.02819,Stanford):七款主流检测器对 91 篇非母语 TOEFL 作文平均误报 61.22%;一行改写提示(”Elevate the provided text by employing literary language”)把检出率从 100% 压到 13% [一手逐字]。
- OpenAI AI Classifier 下线(2023-07-20):官方自述「correctly identifying only 26 percent of AI-written text as ‘likely AI-written’ and incorrectly labeling human-written works 9 percent of the time」,下线原因是「low rate of accuracy」[一手转述官方公告,经 Ars Technica 全文转载]。
- DeGenTWeb(2026-04):
> “when aiming to minimize the chances of falsely attributing human-authored content to LLMs, we find that detectors of LLM-generated text perform much worse than advertised.” [一手逐字]
> “We also show that continuing to accurately identify such sites appears challenging given the capabilities of the latest LLMs.” [一手逐字] - Graphite 的自校准(2026 版):三检测器在 2020–2022 人类文章上的误报率 1.355%–1.844% [一手逐字]——与 Liang/DeGenTWeb 的悲观结论并置,同口径方法学分歧,本篇不裁谁对,只登记。
这一节的小结:「AI 内容占比 X%」的每一个 X 都是一把自带不确定度的尺量出来的——而检测器不确定度在多个独立研究中足以翻转结论。互联网账的事实层能承重的是方向(AI 内容占比在上升),不能承重的是精确刻度(具体百分比随检测器与口径差一个数量级)。
6.3 增量补遗(2026-08-03):人类增量、机器翻译基线、语料管线与水印
〔本节为 2026-08-03 增量,材料来自增量方生态层捆+主笔亲核〕
- 人类增量侧的最硬证据:del Rio-Chanona, Laurentsyeva & Wachs(PNAS Nexus 3(9):pgae400,2024-09,arXiv 版全文)——Stack Overflow 全量 5,800 万帖(2008–2023-06)、以俄语/中文/数学论坛为对照的双重差分:
“A difference-in-differences model estimates a 16% decrease in weekly posts on Stack Overflow. This effect increases in magnitude over time, and is larger for posts related to the most widely used programming languages.” [摘要逐字]
正文:效应「By the end of April 2023, the estimated effect stabilizes at around 25%」,发帖绝对量「falling from around 60,000 posts to 40,000 within six months」;且「Posts made after ChatGPT get similar voting scores than before, suggesting that ChatGPT is not merely displacing duplicate or low-quality content」[一手逐字]。作者对训练数据含义的自述(Discussion):
“our results show that the use of LLMs can slow down the creation of new data. … modelers face the real problem of running out of useful data.” [一手逐字]
——「污染」的另一半:不是 AI 内容变多,是人类公开新增在变少。
- LLM 之前的基线:Thompson et al.(ACL Findings 2024,arXiv:2401.05749):
“Multi-way parallel, machine generated content not only dominates the translations in lower resource languages; it also constitutes a large fraction of the total web content in those languages… Our work raises serious concerns about training models such as multilingual large language models on both monolingual and bilingual data scraped from the web.” [摘要逐字]
——机器翻译内容大规模混入 web 语料在 LLM 时代之前已是事实;「污染」叙事必须扣除这条基线。
- 语料管线的另一半:对 DCLM(arXiv:2406.11794)与 Nemotron-CC(arXiv:2412.02595)全文关键词扫描(AI-generated / machine-generated / LLM-generated / model collapse / watermark),全部零命中(与 5.4 的 FineWeb 同型——增量方 2026-08-03 全文 grep 亲核);Nemotron-CC 方向相反——主动掺入:
“We propose a method for transforming English Common Crawl into a 6.3T token long-horizon pretraining dataset, consisting of 4.4T globally deduplicated original tokens and 1.9T synthetically generated tokens.” [一手逐字] “As our data improves, so will the LLMs we train, and these improved LLMs will in turn improve our data as we use them to generate better synthetic data and quality classifications.” [一手逐字]
——主流公开预训练语料管线对爬取语料中的 AI 内容不设防,合成数据是被刻意加入的特性。
- 水印与「训练前过滤」:Kirchenbauer et al.(ICML 2023,arXiv:2301.10226)——绿表偏置水印,「detectable from short spans of tokens (as few as 25 tokens)」(Intro);OPT-1.3B 自家设定检出 98.4%(z=4 阈值)/99.6%(beam search)[一手逐字]——理想条件、无改写攻击。其动机句直指本层:
“the proliferation of synthetic data on the web complicates future dataset creation efforts, as synthetic data is often inferior to human content and must be detected and excluded before model training.” [一手逐字]
SynthID-Text(Nature 634, 818–823,DOI)已生产部署:
“SynthID-Text has been productionized and is currently watermarking responses in Gemini and Gemini Advanced… the first deployment of a generative text watermark at scale, serving millions of users.” [一手逐字]
约 2,000 万条线上 Gemini 反馈 A/B:好评率差 0.01%、差评率差 0.02%,均不显著——部署未损用户体验。但官方自留三条:「generative watermarks require coordination between actors」「the rise of open-source models presents a challenge」[一手逐字];官方博客另承认「SynthID isn’t a silver bullet… its confidence scores can be greatly reduced when an AI-generated text is thoroughly rewritten or translated」。在网页抓取规模上实证「水印过滤训练语料可行」的研究:增量方检索未找到(照实登记)。
- 监管对商业检测器的官方态度:澳大利亚 ACCC《Digital platform services inquiry》interim report 9(PDF):
“there are no foolproof AI content detection systems, and estimates of the amount of this content online vary significantly.” [一手逐字]
——与 6.2 节检测器反身性互证,且出自政府监管机构而非论战任一方。
七、术语与末日账:dementia、curse、MAD、cannibalism 是怎么长出来的
7.1 改名史:从 dementia 到 collapse
- v1(2023-05-27):《Model Dementia: Generated Data Makes Models Forget》,「We call this effect model dementia」[一手逐字]。
- v2(约 2023-06):改题《The Curse of Recursion: Training on Generated Data Makes Models Forget》,摘要改称 Model Collapse [一手逐字]。
- v3(2024-04-14):脚注 1 自述改名原因——dementia trivialise 医学概念、可能冒犯 [一手逐字](全文见 3.3)。
- Nature 版(2024-07-24):标题《AI models collapse when trained on recursively generated data》,摘要新增「indiscriminate」[一手逐字]。
第一层发现:最吓人的两个词(dementia 痴呆、curse 诅咒)都是作者自己先用的,其中一个被作者自己撤回(dementia)。末日修辞不是媒体单方面加的——它从论文的 v1 标题就开始了。
7.2 疾病隐喻:MAD 与疯牛病
MAD 是论文自己写的类比(「making analogy to mad cow disease」[一手逐字],见 2.2)。媒体的转述层把这个类比进一步肉身化:
- ScienceAlert(2024-08-06):”Cannibal AIs Could Risk Digital ‘Mad Cow Disease’ Without Fresh Data”;引 Baraniuk:”One doomsday scenario is that if left uncontrolled for many generations, MAD could poison the data quality and diversity of the entire internet” [二手转述]。
- Nature News 官方新闻稿(d41586-024-02420-7,2024-07-24):
> “Training artificial intelligence(AI) models on AI-generated text quickly leads to the models churning out nonsense, a study has found. This cannibalistic phenomenon, termed model collapse, could halt the improvement of large language models (LLMs) as they run out of human-derived training data and as increasing amounts of AI-generated text pervade the Internet.” [一手逐字] - Nature 封面(631 期 8022 号)梗:中世纪建筑文本第九代变成「a list of jackrabbits」(长耳大野兔)——这个段子此后成为全球头条标配 [一手逐字]。
- data cannibalism 一词的媒体谱系:Analytics India Magazine 2023-07-18「AI is eating itself. … That’s data cannibalism」[一手逐字,媒体] → TechTarget 2025-07 把它固化进技术百科词条(”training AI on AI is called AI cannibalism, or digital cannibalism”)[一手逐字,媒体]。
7.3 作者团队的两副面孔
- 末日侧:Shumailov 本人在学院新闻稿(Christ Church, Oxford,2024-07-24)里说:
> “‘model collapse’ … refers to AI spiralling into the abyss, feeding on its own mistakes and becoming increasingly clueless and repetitive” [一手逐字,学院新闻稿] - 降温侧:一年后通讯作者 Gal 在 CACM 采访里说:
> “In the end, Gal argued, model collapse is an important consideration, but not the matter of imminent disaster that some news coverage has made it out to be.” [一手逐字,CACM 2025-05-15]
同一团队的两位作者,一个说「螺旋坠入深渊」,一个说「不是媒体写的那种迫在眉睫的灾难」——末日叙事的生产者不只是媒体;而作者之间的口径差,本身就是「崩塌的媒体生命周期」的最好样本。
7.4 学术界对叙事扭曲的正式抗议
- Schaeffer 立场文(venue 存疑见诚实空位 11):「model collapse has been warped from a nuanced multifaceted consideration into an oversimplified threat」[一手逐字](全文见 4.3)。
- The Decoder 采访 Schaeffer(2024-07-30):”It presumes model collapse is a real and significant threat under current best practices. Based on the evidence I’ve seen, it isn’t.” [一手转述]
- The Register(2024-05-09,早于 Nature 发表的争论报道):Gerstgrasser 戏仿末日叙事(”It’s like a virus that could infect the entire AI ecosystem! … But don’t panic just yet”),Shumailov 回应「In principle it does not really invalidate anything we showed.」[一手转述]
文献学更正登记:本篇任务书初稿中的三处记忆线索经实测为伪——「Tartaglione et al. 2023 崩塌论文」(arXiv 作者页全目录无此文);「Koivisto et al. 2025《GPT-generated text in Common Crawl》」(全部公开索引查无此文);「2024 年某研究者用 coprolite(粪化石)形容 AI 生成数据」(仅 2026 年自媒体样本,无学术出处)。另:coprolite 一词有 2026-02 Medium 文章使用(”the forensic analysis of AI training data reveals the health — and the deep, festering sickness — of our own”)[一手逐字,自媒体,不承重]。
7.5 增量补遗(2026-08-03):标题编辑改动实物、三级政策链与 meme 考古
〔本节为 2026-08-03 增量,材料来自增量方叙事考古捆+主笔亲核〕
- 标题编辑改动的实物组(限定词在标题层被丢弃或加重,部分后来被编辑改动留下痕迹):
- TechCrunch(2024-07-24,链接):标题「’Model collapse’: Scientists warn against letting AI eat its own tail」丢限定,正文却逐字保住 Nature 限定句「We discover that indiscriminately learning from data produced by other models causes ‘model collapse’…」[一手逐字]——标题与正文限定词分离的典型样本。
- Forbes(2024-08-26,链接):URL slug 残留「is-ai-quietly-killing-itself」,现标题已弱化为「Is AI Quietly Sabotaging Itself—And The Internet?」——标题被编辑降级过的实物证据(增量方亲核:slug 与现标题并存于同一页)。
- The Register(2024-07-25):论坛存档原标题「AI models face collapse if they overdose on their own output」,现页标题「Recursive training leads to nonsense, study finds」——发表后改题 [一手逐字,两版标题]。
- Live Science(2024-08-09):URL slug「could-spiral-into-unintelligible-nonsense」与现标题「could break down and regurgitate…」不一致——标题后加重过 [一手逐字,slug 与标题]。
- 一句媒体话的三级政策链:Scientific American(2023-07-28,链接)写:
“It suggests that a training diet of AI-generated text, even in small quantities, eventually becomes ‘poisonous’ to the model being trained.” [一手逐字]
——「even in small quantities」是对论文的加码。该句经 News Media Alliance 2023-10-30 意见书进入美国版权局档案,最终出现在美国版权局《Copyright and Artificial Intelligence, Part 3: Generative AI Training》(2025-05,PDF):
“However, use of synthetic data may lead to a phenomenon called model collapse where ‘the outputs quickly start denigrating into nonsense.’ See Authors Guild Initial Comments at 14; Illia Shumailov…” [一手逐字]
——两处转录细节:引号里的「denigrating into nonsense」出自 Authors Guild 意见书而非论文原文(论文无此句),版权局把它与论文并置,制造出「这是论文原话」的印象;且报告把作者名拼成「Illia」。媒体加码句 → 利益方意见书 → 监管报告,三级各丢一层限定。
- meme 考古(「模型吃自己会出事」先于论文的表述):
- Ted Chiang《ChatGPT Is a Blurry JPEG of the Web》(New Yorker,2023-02-09)——Ross Anderson 在博客明说这是他们注意到的先行表述(见下条)。
- Habsburg AI:Jathan Sadowski 的 X 推文(2023-02-13,早于 Shumailov v1 三个多月):
> “I coined a term on @machinekillspod that I feel like needs its own essay: Habsburg AI – a system that is so heavily trained on the outputs of other generative AI’s that it becomes an inbred mutant, likely with exaggerated, grotesque features. It joins the lineage of Potemkin AI.” [一手逐字,经存档页与 The Conversation 内嵌推文双源]
——「needs its own essay」而那篇独立 essay 从未写;该词经 Cory Doctorow《The Coprophagic AI crisis》(2024-03-14)放大,最终进入 Sadowski 专著《The Mechanic and the Luddite》(UC Press,2025)第五章。 - Ross Anderson 博客(2023-06-06,Wayback 存档;原站域名现已被劫持为赌博站,内容以存档为准)——改名内幕与作者限定词的一手出处:
> “This does not mean that LLMs have no uses. As one example, we originally called the effect model dementia, but decided to rename it after objections from a colleague whose father had suffered dementia. We couldn’t think of a replacement until we asked Bard, which suggested five titles, of which we went for The Curse of Recursion. … LLMs are like fire – a useful tool, but one that pollutes the environment.” [一手逐字]
——论文标题由 Google Bard 起名:一个生成式 AI 给「生成式 AI 吃自己会退化」的论文命名,是本篇最反讽的一个脚注。
- 作者自己的限定词(与 7.3 互补):Shumailov 给 VentureBeat 的电邮(2023-06-12,Wayback):
“Note that this does not mean that improbable data should be oversampled, but rather that it should be appropriately represented. As progress drives you to retrain your models, make sure to include old data as well as new.” [一手逐字]
给 Forbes 的电邮(2024-08-26):
“At first, it affects minority data—data that is badly represented. It then affects diversity of the outputs and the variance reduces. Sometimes, you observe small improvement for the majority data, that hides away the degradation in performance on minority data.” [一手逐字]
——「多数数据上的小改善会掩盖少数数据上的退化」是崩塌最隐蔽的形态,与 4.3 的「confidently wrong」同族。
八、反虚无账:崩塌研究不是末日预言,是第一性原理警示
「全是夸大、崩塌不会发生」同样不成立。崩塌研究的价值有四个独立支柱:
- 它是数据循环的第一性原理。Nature 的定理(替换式递归→方差收缩)与「data about human interactions with LLMs will be increasingly valuable」的收尾句([一手逐字]),为「训练数据回路」提供了第一个精确的数学框架——这不是末日预言,是设计约束。
- 它催生了配比科学。EMNLP 2025 的 ~30% 收敛、「崩塌模式出现在教科书式纯合成」的预测(5.3),Kazdan 的「真实数据稀缺时存在合成数据最优量」(5.3)——崩塌理论直接转化为工程配比依据。
- 它催生了验证框架。《Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification》(OpenReview 2024):
“we investigate the use of verification on synthesized data to prevent model collapse” [一手逐字]
以及 KITE(arXiv:2607.17043,2026-07,崩塌可表现为「能力极化」——强化强者、恶化弱者,而非均匀退化)、REFED(EACL 2026,合成数据有质量天花板——「models trained on the data cannot outperform the LLM generating it」)、EvoSyn(ACL 2026,一致性评估器)——合成数据研究正在变成「验证器设计」学科 [一手逐字]。
- 生态视角:《Data Pollination》(ACL 2026,320 个模型种群实验):
“We term this process data pollination, the unintentional circulation of synthetic model outputs through shared online platforms and web-scale training corpora” … “ecological diversity functions as a fundamental resilience mechanism that safeguards the ecosystem against collapse” [一手逐字]
——崩塌研究把「数据循环」从一个比喻变成了一个可建模、可实验、可设计的对象。
反向红线:崩塌「被夸大」不等于「没发生」——替换式递归的数学是真的(第二节)、Dohmatob 的 1% 强崩塌是真的(4.2)、FineWeb 16% 检出是真测量(5.4);批评者自己也承认「several prominent collapse scenarios are readily avoidable」(Schaeffer [一手逐字])——「可避免」的前提是知道条件,这正是本篇第三节整节的工程量。
九、边界账:版权与合规接口(只登记)
- NYT v. OpenAI:第三修正起诉状(2026-06-25,72 页)全文
synthetic出现 31 处,全部是「synthetic search results / synthetic output」语境——无一处指合成数据作训练数据。至 2026-06 本案训练数据合成化未成争点 [一手逐字,起诉状全文计数]。只登记,不作诉讼预测。 - 欧盟 AI 法案:第 50 条(2024-06-13 通过、2024-07-12 公报、2026-08-02 起适用)是对输出端的披露义务:
“2. Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.” [一手逐字,官方公报] “4. …Deployers of an AI system that generates or manipulates text which is published with the purpose of informing the public on matters of public interest shall disclose that the text has been artificially generated or manipulated.” [一手逐字]
训练数据侧在 53(1)(c)(版权政策)与 53(1)(d)(训练内容摘要公开义务)——第 50 条管「输出可识别」,不管「训练数据配比」。
- C2PA:2.4 规范已把「Training & Data Mining」断言从主规范移除、移交 Creator Assertions WG(cawg.io/training-and-data-mining/1.1,allowed/notAllowed/constrained 三值)——训练数据溯源正从「内容凭证」走向独立工作组 [一手逐字,官方规范页]。
- 政策话语中的崩塌引用谱系〔2026-08-03 增量登记,六件〕:崩塌概念进入政策文件的引用忠实度光谱——
- 最忠实:国际清算银行《Annual Economic Report 2024》第三章脚注 31:
> “Synthetic data are unlikely to help. … continued reliance on these data by AI models diminishes the information coming from the tails of the distribution (ie rare but highly consequential events)… See Shumailov et al (2023).” [一手逐字]
——精确落在「尾部信息丢失」,且与宏观审慎政策(稀有高后果事件)的语境对齐。 - 引无限定版:EPIC 给 NIST 的意见书(2024-02-02,PDF)引 arXiv 版定义(故无 indiscriminately)。
- 平直转述:Bruegel 政策简报(2024-01-25,链接):”Using synthetic training data may cause model ‘collapse’, meaning a dramatic drop in model output quality.”
- 缓冲派引用:法国文化部 CSPLA 报告(PDF)用 Gerstgrasser/Kazdan 的 accumulation 框架论证「catastrophic collapse is unlikely, at least after a few generations」——政策文件里少见的缓冲派。
- 委员会研究:欧洲议会 JURI《The Development of Generative Artificial Intelligence from a Copyright Perspective》(2025-05-12,PDF)引 Alemohammad 与 Shumailov 二文作「overreliance… may result in declining output quality and a phenomenon known as model collapse」。
- 产业界科普:IBM Think《What Is Model Collapse?》(2024-10-14,链接);OpenAI/Google DeepMind/Meta/Anthropic 官方站点未检索到使用该概念(增量方检索,缺席不等于不存在)。
边界账只登记三个事实:法庭文书、法条、标准规范各自对「合成数据」的处理——它们都还不构成对「训练数据合成占比」的任何监管,全部是输出侧或版权侧。
十、母题收口:三个问题与一把旋钮
把全篇压成三问一旋钮:
问一:报的是什么? 崩塌是「多代递归」的定理,不是「训练数据里有 AI 内容」的定理。互联网占比(第六节)与崩塌定理(第二节)之间隔着一个条件——递归采样。媒体把「互联网 AI 内容占比 ~50%」和「递归训练必然崩塌」缝在一起时,缝的不是证据,是省略号。
问二:谁在什么设定下说的? Nature 说 inevitable 时带了三个条件(indiscriminate/recursive/ideal conditions 下统计误差即可致崩);Gerstgrasser 说不 inevitable 时也带了条件(accumulate);Dohmatob 说 1% 够时带了条件(回归、p₂ 不趋于零)。两边都没有说过无条件的话——无条件的那句是头条标题。
问三:工程在干什么? 工程从来不是「用合成数据」或「不用合成数据」的二选一,而是五个旋钮:类型(改写 vs 教科书式)、配比(~20–40%)、规模(大模型更韧)、验证(可验证任务收益最大)、新鲜数据(保留比例)。phi-4 的消融是这五个旋钮最诚实的自白:同段两句,一句「多轮次合成优于新网页」,一句「纯合成掉分加幻觉」。
旋钮在谁手里? 旋钮在造数据的人手里。模型不会拧自己的旋钮——「模型死于自己的输出」是拟人化的自噬跳;准确的说法是「配比规则写错时,多代训练会退化」,而配比规则是人写的。
灵魂句:崩塌是真的、合成数据有用也是真的——真的不是「AI 即将死于自己的输出」和「合成数据是免费的无限燃料」这两层被声称的定论;把我们的配比规则读成模型的宿命,是把一道工程选择题读成了一句死刑宣判。定理的每个限定词都在原文里写着;丢限定词的那一步,发生在从论文到头条的路上,而这条路是媒体、读者和研究者合修的。
十一、诚实空位
- Ilya Sutskever「数据是化石燃料/peak data」:无官方逐字稿,仅两份社区转录(内容一致),本篇不承重为逐字引用,只登记为转述。
- 「2023 年学者估计 AI 内容占比约 2–3%」:未找到一手出处;最接近的一手是 Copyleaks 2024-03 的 1.57%(含 AI 页面口径)。
- Copyleaks 2024 研究官方发布页:已下架;数字经两源交叉核对(TechNewsWorld + 官方博客),标 ◐。
- 「崩塌 vs 幻觉」的学术区分:未找到专门论文作定义层区分,只登记媒体层接链(ScienceNews 转述 Shumailov)。
- Gerstgrasser 组 OpenReview 页面:被 Cloudflare 拦截,NeurIPS 2024 收录状态以 dblp(仅 CoRR)为准;会议版以 ICML 2025 收录为准。〔2026-08-03 增量方亲核修订〕OpenReview API 检索显示 2404.01413 另有 COLM 会议版记录(venue 字段逐字 “COLM”)——即该文有 COLM 与 ICML 2025《Collapse or Thrive?》两个会议版,「会议版以 ICML 2025 收录为准」一句据此修订。
- diffusion/dreambooth 迭代专项的 2025-2026 新论文:仅找到 Chain of Diffusion(2407.17493)与一篇仅标题收录(2602.19033),未取回摘要。
- 各实验室合成配比官方数字:GPT-5 系统卡 0 次、DeepSeek-V3 全文 0 次、Llama 4 未披露——「配比不披露」本身是系统性事实,故第五节的配比证据全部来自学术论文(phi-4/EMNLP/Kazdan)而非厂商。
- coprolite 词源:无学术出处,仅自媒体使用,未承重。
- 「占比→模型质量下降」的因果链:本库未找到「训练语料 AI 占比上升→某模型质量下降」的直接实证(DeGenTWeb 的 2604.26965 甚至报告「未发现事实准确性或文体多样性恶化的统计显著证据」[一手逐字,摘要级]);按纪律如实写「本篇不建立这条链」。
- 本篇自己的完工判据:写在这——若出现下列任一证据,本篇第三/四节的结论需重审:(a) 某前沿模型被证实实际发生了替换式多代递归训练并观测到崩塌;(b) 「互联网占比→模型退化」的直接因果实证;(c) 崩塌定义在学界收敛为单一标准(第八节 Schaeffer 八定义的争议被解决)。否则本篇裁决视为有效。
- Schaeffer 立场文的 venue〔2026-08-03 增量方登记〕:本篇 4.3 与 7.4 两处标「ICML 2026」;增量方经 OpenReview API 检索仅见「CoRR 2025」与「Submitted to NeurIPS 2025 Position Paper Track」两条公开记录,未见 ICML 2026 收录记录——按纪律登记存疑,不断言错误,引用者请以 arXiv:2503.03150 页为准。
来源清单
- Shumailov et al. 2024, Nature 631, 755–759《AI models collapse when trained on recursively generated data》:https://www.nature.com/articles/s41586-024-07566-y [一手逐字]( 摘要+正文 Theorem 3.1/Definition 2.1/三误差/实验/引言)
- arXiv v3《The Curse of Recursion》(2024-04-14):https://arxiv.org/abs/2305.17493 [一手逐字]( 摘要/1D 递推/重复惩罚排除)
- arXiv v1《Model Dementia》(2023-05-27):https://arxiv.org/abs/2305.17493v1 [一手逐字]( 原始标题与摘要)
- Alemohammad et al. 2023《Self-Consuming Generative Models Go MAD》:https://arxiv.org/abs/2307.01850 [一手逐字]( MAD 定义/三种循环/结论句)
- Hataya, Bao & Arai, ICCV 2023《Will Large-scale Generative Models Corrupt Future Datasets?》:https://arxiv.org/abs/2211.08095 [一手逐字]( 摘要)
- Bohacek & Farid 2023《Silent Self-Generated Data…》:https://arxiv.org/abs/2311.12202 [一手逐字]( 摘要)
- Bertrand et al., ICLR 2024:https://arxiv.org/abs/2310.00429 [一手逐字]( 摘要)
- Briesch et al. 2023:https://arxiv.org/abs/2311.16822 [一手逐字]( 摘要)
- Gerstgrasser et al. 2024《Is Model Collapse Inevitable?》:https://arxiv.org/abs/2404.01413 [一手逐字]( 摘要/Theorem 2/Discussion/保留条款/术语四义)
- Gerstgrasser et al., ICML 2025《Collapse or Thrive?》:https://arxiv.org/abs/2410.16713 [一手逐字]( 摘要三种工作流)
- Dohmatob et al. 2024《Model Collapse Demystified》:https://arxiv.org/abs/2402.07712 [一手逐字]( 摘要/Thm 4.9/Remark 4.2)
- Dohmatob et al. 2024《Strong Model Collapse》:https://arxiv.org/abs/2410.04840 [一手逐字]( 摘要/Result #1)
- Barzilai & Shamir 2025《When Models Don’t Collapse》:https://arxiv.org/abs/2505.19046 [一手逐字]( 摘要)
- Schaeffer et al.《Position: Model Collapse Does Not Mean What You Think》:https://arxiv.org/abs/2503.03150 [一手逐字]( 摘要)
- Knowledge Collapse in LLMs:https://arxiv.org/abs/2509.04796 [一手逐字]( 摘要)
- Anti-Ouroboros Effect:https://arxiv.org/abs/2509.10509 [一手逐字]( 摘要)
- Machine-generated text detection prevents model collapse(EMNLP 2025):https://aclanthology.org/2025.emnlp-main.1506.pdf [一手逐字]( 摘要)
- Escaping Model Collapse via Synthetic Data Verification:https://arxiv.org/abs/2510.16657 [一手逐字]( 摘要)
- ChatGPT 多样性纵向观测:https://arxiv.org/abs/2603.12683 [一手逐字]( 摘要)
- SIGMA: Scalable Spectral Insights:https://arxiv.org/abs/2601.03385 [一手逐字]( 摘要)
- Self-Instruct:https://arxiv.org/abs/2212.10560 [一手逐字]( 摘要)
- WizardLM:https://arxiv.org/abs/2304.12244 [一手逐字]( 摘要)
- Orca:https://arxiv.org/abs/2306.02707 [一手逐字]( 摘要)
- Orca 2:https://arxiv.org/abs/2311.11045 [一手逐字]( 摘要)
- phi-1:https://arxiv.org/abs/2306.11644 [一手逐字]( 摘要)
- phi-1.5:https://arxiv.org/abs/2309.05463 [一手逐字]( 摘要)
- phi-2 模型卡:https://huggingface.co/microsoft/phi-2 [一手逐字]( 模型卡)
- phi-4:https://arxiv.org/abs/2412.08905 [一手逐字]( 摘要/引言 bulk 句/§3.1 Table 5/消融两句)
- Cosmopedia HF 博客:https://huggingface.co/blog/cosmopedia [一手逐字]( 博客)
- Llama 4 官方博客:https://ai.meta.com/blog/llama-4-multimodal-intelligence/ [一手逐字]( 博客正文,经 Exa 缓存取回)
- Llama 4 模型卡:https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md [一手逐字]( 模型卡)
- Self-Rewarding LMs:https://arxiv.org/abs/2401.10020 [一手逐字]( 摘要)
- DeepSeek-R1:https://arxiv.org/abs/2501.12948 [一手逐字]( 摘要/正文 AIME 15.6→77.9/蒸馏 80 万样本)
- Altman CNBC 采访转录(2025-08-08):https://www.cnbc.com/2025/08/08/first-on-cnbc-transcript-openai-sam-altman-speaks-with-cnbcs-squawk-box-today.html [一手逐字]( 转录稿,页面自标 unofficial)
- GPT-5 系统卡:https://cdn.openai.com/gpt-5-system-card.pdf [一手逐字]( 全文 synthetic 0 次)
- Claude Sonnet 5 系统卡:https://www-cdn.anthropic.com/283ef97c476cf442c91d9a37d5b214242a55bb92/Claude%20Sonnet%205%20System%20Card.pdf [一手逐字]( 训练数据段)
- 扎克伯格 Dwarkesh 采访(2024-04-18):https://www.dwarkesh.com/p/mark-zuckerberg [一手逐字]( 文字稿)
- EMNLP 2025《Demystifying Synthetic Data in LLM Pre-training》:https://arxiv.org/abs/2510.01631 [一手逐字]( 摘要)
- Synthetic Data Proportions:https://arxiv.org/abs/2510.05133 [一手逐字]( 摘要)
- Kazdan et al., ICML 2025:https://proceedings.mlr.press/v267/kazdan25a.html [一手逐字]( 摘要)
- Barbaro et al. 2025(HAL):https://hal.science/hal-05458661v1/document [一手逐字]( 摘要)
- MetaMath:https://arxiv.org/abs/2309.12284 [一手逐字]( 摘要)
- WizardCoder:https://arxiv.org/abs/2306.08568 [一手逐字]( 摘要)
- DeepSeek-Prover:https://arxiv.org/abs/2405.14333 [一手逐字]( 摘要)
- Ahrefs 2025:https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated/ [一手逐字]( 74.2%/2.5%/71.7%)
- Graphite 2026:https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do [一手逐字]( 49.6/50.9/49.9/35.9/55.4k/检测器自校准)
- Graphite 2025-10 版:https://graphite.io/five-percent/research/more-articles-are-now-created-by-ai-than-humans [一手逐字]( 3.3pp 差异)
- Spennemann 2025:https://arxiv.org/abs/2504.08755 [一手逐字]( 摘要 ≥30% 近 40%)
- Dolezal et al. 2026:https://arxiv.org/abs/2604.26965 [一手逐字]( 摘要 35%)
- DeGenTWeb:https://arxiv.org/abs/2605.00087 [一手逐字]( 摘要/2.1%→29.4%/检测器 worse than advertised)
- NBER WP 34777(Reimers & Waldfogel):https://www.nber.org/papers/w34777 [一手逐字]( 摘要/正文 30/45/60%)
- Copyleaks 2024(两源交叉):https://www.technewsworld.com/story/copyleaks-study-finds-explosive-growth-of-ai-content-on-the-web-179161.html [◐]( 1.57%/8,362%)
- 维基百科 AI 内容(Brooks et al. 2024):https://arxiv.org/abs/2410.08044 [一手逐字]( 摘要 4.36%)
- Liang et al. 2024(学术论文 LLM 修改):https://arxiv.org/abs/2404.01268 [一手逐字]( 摘要 CS 17.5%)
- NewsGuard AI Tracking Center:https://www.newsguardtech.com/special-reports/ai-tracking-center/ [一手逐字]( 2026-06 3,749 个)
- Liang et al. 2023(检测器误报):https://arxiv.org/abs/2304.02819 [一手逐字]( 摘要 61.22%/13%)
- OpenAI AI Classifier 下线(Ars 转载官方公告):https://arstechnica.com/information-technology/2023/07/openai-discontinues-its-ai-writing-detector-due-to-low-rate-of-accuracy/ [◐]( 官方公告转述 26%/9%)
- 欧盟 AI 法案(OJ L 2024/1689)第 50 条:https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32024R1689 [一手逐字]( 50(2)/50(4))
- AI 法案第 50 条适用日期(欧盟官方 FAQ):https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act [一手逐字]( 2026-08-02)
- C2PA 2.4 AI/ML 规范:https://spec.c2pa.org/specifications/specifications/2.4/ai-ml/ai_ml.html [一手逐字]( 官方)
- NYT v. OpenAI 第三修正起诉状(2026-06-25):https://cdn.arstechnica.net/wp-content/uploads/2026/06/NYT-v-OpenAI-Third-Amended-Complaint-6-25-26.pdf [一手逐字]( 全文 synthetic 31 处计数)
- Nature News 2024-07-24:https://www.nature.com/articles/d41586-024-02420-7 [一手逐字]( cannibalistic 句)
- Nature 631 期封面页:https://www.nature.com/nature/volumes/631/issues/8022 [一手逐字]( jackrabbits)
- Shumailov 学院新闻稿(Christ Church 2024-07-24):https://chch.ox.ac.uk/news/could-machine-learning-models-cause-their-own-collapse [一手逐字]( spiralling into the abyss)
- Gal CACM 采访(2025-05-15):https://cacm.acm.org/news/the-collapse-of-gpt/ [一手逐字]( not the matter of imminent disaster)
- ScienceAlert 2024-08-06:https://www.sciencealert.com/cannibal-ais-could-risk-digital-mad-cow-disease-without-fresh-data [二手]( 疯牛病转述)
- Analytics India Magazine 2023-07-18:https://analyticsindiamag.com/ai-trends/the-dark-consequence-of-ais-data-cannibalism [一手逐字,媒体]( data cannibalism)
- TechTarget AI cannibalism 词条(2025-07):https://www.techtarget.com/whatis/feature/AI-cannibalism-explained [一手逐字,媒体]( 词条定义)
- The Decoder 采访 Schaeffer(2024-07-30):https://the-decoder.com/ai-data-isnt-destroying-ai-models-after-all-researchers-say/ [一手转述]()
- The Register(2024-05-09):https://theregister.com/software/2024/05/09/experts-divided-over-training-ai-with-more-data-from-ai/1202288/ [一手转述]( 戏仿末日叙事)
- Beyond Model Collapse(OpenReview 2024):https://openreview.net/forum?id=MQXrTMonT1 [一手逐字]( 摘要 verification)
- KITE(Kazdan 2026):https://arxiv.org/abs/2607.17043 [一手逐字]( 摘要能力极化)
- Data Pollination(ACL 2026):https://aclanthology.org/2026.acl-long.1229.pdf [一手逐字]( 摘要)
- Villalobos et al.《Will We Run Out of Data?》v2:https://arxiv.org/abs/2211.04325 [一手逐字]( 摘要/正文中位 2028)
- Epoch《AI in 2030》:https://epoch.ai/files/AI_2030.pdf [一手逐字]( unlikely to run out of data/exhausted before 2027/2.7x/合成+多模态)
- 多语言合成数据退化(ACL 2026):https://aclanthology.org/2026.acl-long.1002.pdf [一手逐字]( 摘要)
- 低资源语言(Welsh −11.53%):https://arxiv.org/abs/2506.12158 [一手逐字]( 摘要)
- 维基整体 LLM 影响 1-2%:https://arxiv.org/abs/2503.02879 [一手逐字]( 摘要)
以下 32 条为 2026-08-03 增量(Kimi 六捆调研+主笔亲核):
- Seddik et al. 2024《How Bad is Training on Synthetic Data?》:https://arxiv.org/abs/2404.05090 [一手逐字]( 摘要/Theorem 1/Proposition 1)
- Ferbach et al., NeurIPS 2024《Curated Data Provably Optimize Human Preferences》:https://arxiv.org/abs/2407.09499 [一手逐字]( 摘要/Theorem 2.1)
- Dohmatob et al., ICML 2024《A Tale of Tails》:https://arxiv.org/abs/2402.07043 [一手逐字]( Theorem 2.1/3.2)
- Kovač et al., EMNLP 2025 Oral《Recursive Training Loops in LLMs》:https://arxiv.org/abs/2504.03814 [一手逐字]( r=1/16…1/2 剂量曲线)
- Qiao et al. 2026《When Sample Selection Bias Precipitates Model Collapse》:https://arxiv.org/abs/2606.13732 [摘要逐字]( 低资源验证孤岛)
- Roe et al. 2026《Iterative Finetuning is Mostly Idempotent》:https://arxiv.org/abs/2605.01130 [摘要逐字]( 特质衰减)
- 《When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs》2026:https://arxiv.org/abs/2606.28438 [一手逐字]( 自评门控退化)
- Trinh et al., Nature 625, 476–482《AlphaGeometry》:https://doi.org/10.1038/s41586-023-06747-5 [一手逐字]( Main 节 100 million 句)
- NVIDIA Nemotron-4 340B 技术报告:https://arxiv.org/abs/2406.11704 [摘要逐字]( 对齐 98% 合成)
- Magpie:https://arxiv.org/abs/2406.08464 [摘要逐字]( 300K vs 10M)
- SmolLM2(含 SmolTalk/Magpie-Ultra):https://arxiv.org/abs/2502.02737 [一手逐字]( §4.5/§5.1)
- RoentGen(胸片 latent diffusion):https://arxiv.org/abs/2211.12737 [摘要逐字]( 混合 +5%/纯合成 +3%)
- del Rio-Chanona et al., PNAS Nexus 3(9):pgae400(Stack Overflow):https://arxiv.org/abs/2307.07367 [一手逐字]( 摘要 −16%/Discussion「running out of useful data」)
- Thompson et al., ACL Findings 2024(机器翻译基线):https://arxiv.org/abs/2401.05749 [摘要逐字]
- DCLM:https://arxiv.org/abs/2406.11794 [增量方全文 grep 零命中亲核]
- Nemotron-CC:https://arxiv.org/abs/2412.02595 [一手逐字]( 4.4T+1.9T/飞轮句)
- Kirchenbauer et al., ICML 2023(文本水印):https://arxiv.org/abs/2301.10226 [一手逐字]( Intro 25 tokens/§4.1 98.4%)
- Dathathri et al., Nature 634, 818–823(SynthID-Text):https://doi.org/10.1038/s41586-024-08025-4 [一手逐字]( 部署句/Limitations)
- 澳大利亚 ACCC《Digital platform services inquiry》interim report 9:https://www.accc.gov.au/system/files/digital-platform-services-inquiry-report9_0.pdf [一手逐字]( no foolproof 句)
- TechCrunch 2024-07-24:https://techcrunch.com/2024/07/24/model-collapse-scientists-warn-against-letting-ai-eat-its-own-tail/ [一手逐字]( 标题丢限定/正文保留)
- Forbes 2024-08-26:https://www.forbes.com/sites/torconstantino/2024/08/26/is-ai-quietly-killing-itself-and-the-internet/ [一手逐字]( slug killing/现标题 Sabotaging)
- Scientific American 2023-07-28:https://www.scientificamerican.com/article/ai-generated-data-can-poison-future-ai-models/ [一手逐字]( even in small quantities)
- 美国版权局《Copyright and AI, Part 3》2025-05:https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf [一手逐字]( Authors Guild 引句/Illia 拼写)
- Ross Anderson 博客 2023-06-06(Wayback;原站域名已被劫持):http://web.archive.org/web/2023/https://www.lightbluetouchpaper.org/2023/06/06/will-gpt-models-choke-on-their-own-exhaust/ [一手逐字]( 改名内幕/Bard 起名/like fire)
- VentureBeat 2023-06-12(Wayback):http://web.archive.org/web/2023id_/https://venturebeat.com/ai/the-ai-feedback-loop-researchers-warn-of-model-collapse-as-ai-trains-on-ai-generated-content/ [一手逐字]( Shumailov 电邮限定词)
- Sadowski「Habsburg AI」推文 2023-02-13(存档双源):https://markcarrigan.net/2023/04/10/habsburg-ai-a-system-that-is-so-heavily-trained-on-the-outputs-of-other-generative-ais-that-it-becomes-an-inbred-mutant/ [一手逐字]
- BIS《Annual Economic Report 2024》第三章脚注 31:https://www.bis.org/publ/arpdf/ar2024e3.htm [一手逐字]
- EPIC 给 NIST 意见书 2024-02-02:https://epic.org/wp-content/uploads/2024/02/EPIC-Comment-on-NIST-AI-Executive-Order-Mandates-RFI-02.02.24.pdf [一手逐字]
- Bruegel 政策简报 2024-01-25:https://www.bruegel.org/policy-brief/why-artificial-intelligence-creating-fundamental-challenges-competition-policy [一手逐字]
- 法国文化部 CSPLA 报告:https://www.culture.gouv.fr/Media/medias-creation-rapide/cspla-ai-cultural-content-remuneration-economics-component-english.pdf [一手逐字]
- 欧洲议会 JURI 研究 2025-05-12:https://www.europarl.europa.eu/meetdocs/2024_2029/plmrep/COMMITTEES/JURI/DV/2025/05-12/2025.05.12_item6_Study_GenAIfromacopyrightperspective_EN.pdf [一手逐字]
- IBM Think《What Is Model Collapse?》2024-10-14:https://www.ibm.com/think/topics/model-collapse [登记]
来源分级统计:110 条来源(原 78 条+2026-08-03 增量 32 条),[一手逐字] 96 条、[一手转述] 3 条、[◐ 双源交叉] 2 条、[二手转述] 1 条、[摘要逐字] 5 条、[登记] 1 条、未取回登记 6 处(见诚实空位)。外链实测:本篇正文唯一外部链接经逐一 curl 验证,全部可访问(个别出版社/防爬站点按 runbook 改走官方替代或 Wayback 原件,详见生成器注释)。
调研纪律:六捆并行一手调研(守真锚机制账/条件账/工程账/互联网账/术语末日账/反向整体账)+主笔亲核 16 处承重引用(Nature 摘要与 Theorem 3.1/Gerstgrasser 摘要与 Theorem 2 与 Discussion/phi-4 消融两句与 bulk 句/R1 摘要与 AIME 数字/Epoch AI in 2030 逐字/Graphite 全部关键数字/AI 法案第 50 条/DeGenTWeb 摘要/v1 Model Dementia 标题/Schaeffer 八定义/Altman 转录稿)全部逐字命中。任务书级纠错七处:Gerstgrasser 非 NeurIPS 2024(dblp 仅 CoRR,会议版 ICML 2025);scaling 篇 phi-4 消融标注 §4 实为 §3.1;Cosmopedia 无 arXiv 论文(2402.10200 是《CoT Without Prompting》);Tartaglione 2023 崩塌论文查无此文;Koivisto 2025 Common Crawl 检测查无此文;coprolite 无 2024 学术出处;Nature News URL d41586-024-02359-1 已失效(正确 02420-7)。去重实测:模型崩塌 中文零命中,model collapse/synthetic data 全库 18 处集中在 scaling 篇,本篇增量全部在机制/条件/互联网/术语/工程五层,与 scaling 篇分界写死。
2026-08-03 增量合并记:本篇 08-02 落档同日,另一流程(Kimi)按头儿 08-01 点定独立完成同主题调研(其查重时本篇尚未落档)——六捆并行一手调研(理论地基/实验谱系/工程实证/生态层测量/叙事考古版本学/批判与最新发展)+主笔亲核约 60 组承重引用(arXiv 摘要 24 篇、Europe PMC 全文 3 篇、Crossref 3 篇、官方文件与媒体 15 处、语料论文全文 grep 6 篇、OpenReview API 4 处),全部逐字命中。经头儿裁定不做重复篇,增量合并为本篇 4.3.1/5.5/6.3/7.5 四节、第九节第 4 条、诚实空位第 5 条修订与第 11 条新增、来源 79–110 号。增量方独立复现两处文献学更正(Gerstgrasser 非 NeurIPS 2024——并另补其 COLM 2024 会议版记录;Barzilai & Shamir 编号 2505.19046 验证一致),新增 Strong Model Collapse arXiv「1%」vs ICLR 会议版「1 per 1000」版本差与 Schaeffer venue 存疑登记。