目录
- 零、一句话裁决
- 一、本篇测什么:三层承重与六件可判定的事
- 二、尺子跳:每个榜声称测什么
- 2.1 GLUE(2018):声称推动「通用语言理解」
- 2.2 SuperGLUE(2019):因为 GLUE 被追平才诞生
- 2.3 MMLU(2020):声称测「预训练获得的知识」
- 2.4 MMLU-Pro(2024):因为 MMLU 饱和才诞生
- 2.5 HumanEval(2021):声称测「功能正确性」
- 2.6 Chatbot Arena(2024):声称用人类偏好替代静态榜
- 2.7 ARC-AGI(2019/2025):声称测「流体智能」——名字自带 AGI
- 2.8 SWE-bench 系列(2023-2026):声称测「真实软件工程」
- 2.9 Open LLM Leaderboard(2023-2025):声称「开放、公平、可复现」
- 2.10 Stanford HELM(2022-):声称「可复现、透明」
- 三、分母跳:测试集与训练集的边界——已证实的泄漏事实账
- 3.1 开山之作:Dodge et al. 2021 的 C4 重叠分析
- 3.2 WIMBD(2023):67/82 干净——但被污染的 15 个里藏着重点
- 3.3 Li et al.(2023/2024):MMLU 29.1% 逐字重叠于 Common Crawl
- 3.4 GPT-4 技术报告(2023):官方承认三件事
- 3.5 Gemini 1.0(2023):官方承认「微调 100 步就能提分」
- 3.6 Claude Opus 4.5 系统卡(2025-11):官方承认「改写题照样漏网」
- 3.7 OpenAI 2026-02-23:SWE-bench Verified 官方宣布「废掉」
- 3.8 其他已证实的案例清单
- 四、检测会计学:各家自查方法能检出什么
- 五、记分卡账:第三方探针——「背题」的直接证据
- 六、排行榜会计学:分数的可比性、解析器、权重与退役
- 七、饱和与反饱和账:SWE-bench 三圈循环
- 八、名号跳:榜分→能力→智能→国家竞争力
- 九、反向账:尺子不完美≠尺子无用
- 十、方法论账、裁决表与诚实空位
- 来源清单(编号)
机制裁决第 116 篇 · 对称双向第 111 篇 · AI/认知谱系 · 全库第 174 篇 本篇是对 AI 评测榜与数据污染的系统审计。先读三句红线:
- 本篇不做「哪个模型好、哪个公司坏」的裁决,不评价任何在世研究者,也不给「该信哪个榜」的产品建议——只审「榜上的分数是测量还是叙事」。
- 「数据污染」与「排行榜无用」是两件不同的事:前者是有实证的工程问题,后者是常见于社交媒体的情绪结论,本篇对称审计。
- 全篇承重句均给出可点击来源;正文中英文逐字引用一律标注「一手逐字」或「摘要逐字」,未取回原文的明确降档。
零、一句话裁决
AI 评测榜是尺子,而「榜分=能力=智能」是升格链。 这把尺子真(它推动过 BERT 时代、被全行业当验收单用、且被反复维修过),但它的刻度一直在被三件事污染:测试集与训练集的边界正在消失(MMLU 29.1% 逐字重叠于 Common Crawl [一手逐字]、GPT-4 官方自认 GSM-8K 训练集进过预训练 [一手逐字])、各家自查的方法各有系统性盲区(GPT-4 抽 3 段 50 字符 [一手逐字],被 ConTAM 判「太严格、必然漏检」[一手逐字])、以及榜分被读成能力、智能与国家竞争力的名号升格(o1「超过人类博士」[一手逐字]、AI Index 以 MMLU 分差排国家 [一手逐字])。而「全是刷分游戏」同样未立:GLUE 驱动了 BERT 时代 [文献较稳]、ETS 用 70 年证明「有作弊≠考试失效」[一手逐字]、SAT 补习几十年只把平均分抬高约 30 分——恰好等于测量误差 [一手逐字]。落点:AI 从来不缺分数,缺的是对分数的效度论证——2026 年最硬的证据是 OpenAI 自己宣布停止报告 SWE-bench Verified 分数,因为「所有前沿模型都见过题目」[一手逐字]。
一、本篇测什么:三层承重与六件可判定的事
本篇审的不是「AI 强不强」,而是「我们凭什么知道 AI 强不强」。拆成六件可判定的事:
- 尺子跳:每个主流评测榜声称测什么(GLUE 声称推动通用 NLU [一手逐字]、MMLU 声称测预训练获得的知识 [一手逐字]、ARC-AGI 声称测流体智能 [一手逐字])——声称与实测之间的距离是第一层承重。
- 分母跳:测试集与训练集的边界在哪(已证实的泄漏案例清单:谁、何时、证据形式、是否官方承认)。
- 检测会计学:各家自查方法(GPT-4 的 3×50 字符、Llama 2 的 token 级 n-gram+统计检验、Anthropic 三代模型卡零污染章节 [一手逐字]、Gemini 的定性弃报)各自能检出什么、检不出什么,有没有独立复算。
- 排行榜会计学:分数可比性(微扰排名移动 8 位 [一手逐字])、解析器换挡洗牌 Top20 [一手逐字]、归一化=改权重 [一手逐字]、饱和与退役账。
- 名号跳:榜分→能力→智能→国家竞争力的升格链(o1 博士 [一手逐字]、AI Index 排国家 [一手逐字]、DeepSeek 与 o1 逐分比差 [一手逐字])。
- 反向红跳:尺子不完美≠尺子无用(GLUE 历史作用、Messick 效度框架 [一手逐字]、ETS 考试安全工程 [一手逐字]、SAT 补习 30 分实证 [一手逐字])。
结构胎记:尺子跳 × 分母跳 × 名号跳,附一条反向红跳——与库内各篇的分界写死如下。
与库内相关篇的分界:
- 与 AGI 篇(2026-08-01)分界:那篇把 ARC-AGI-2 的 85% 门槛与 24.03% 榜首当作「AGI 完工判据」的活体证据;本篇正面审计评测榜本身——ARC-AGI-2 的抗污染设计(pass@2、IDD 校准、changelog「reduce data mining and overfiting」[一手逐字])在本篇是「反饱和机制」的案例,不是判据。
- 与 scaling 篇(2026-07-23)分界:那篇把 benchmark 分数当作「能力轨迹」的证据;本篇审「轨迹的刻度准不准」——同一条 MMLU 曲线,在「分数=能力」与「分数=记忆」之间摇摆。
- 与测量代理篇(2026-06-22)分界:那篇是方法学(proxy 的 Goodhart 化);本篇是它的活体应用——「MMLU 分数」作「AI 能力」的 proxy,「排行榜」作「进展」的 proxy。
- 与过拟合篇(2026-05-05)分界:那篇是 AI 学习冥想场景;本篇的 overfitting 是「基准过拟合」(benchmark overfitting)——一个不同构的靶子。
- 与量子霸权篇(2026-07-09)分界:RCS 是「采样任务验收」(Supremacy≠Advantage);本篇是通用评测榜,但共享「验收单被改规则」的母题。
- 与地震预测篇(2026-08-01)分界:CSEP 前瞻评分 vs 回顾性拟合的落差,与本篇「回顾性 benchmark vs 前瞻性 hidden eval」完全同构——但对象不同(地震预测 vs AI 评测),本篇不重做那篇的统计账。
- 与数字孪生篇(2026-07-26)分界:那篇审「模拟器=你」;本篇审「分数=能力」。共享的教训:测量对象是「解释」不是「工具」(Messick:被评估的是分数解释,不是测试本身 [一手逐字])。
编号落位:机制裁决第 116 篇·对称双向第 111 篇·全库第 174 篇(2026-08-01 全库扫描确认未被占用)。
二、尺子跳:每个榜声称测什么
先清点尺子本身。本库原则:审「声称」必须回到声称的原件。
2.1 GLUE(2018):声称推动「通用语言理解」
GLUE 论文(arXiv:1804.07461,2018-04-20,ICLR 2019)的动机陈述原文:
“If we aspire to develop models with understanding beyond the detection of superficial correspondences between inputs and outputs, then it is critical to develop a more unified model that can learn to execute a range of different linguistic tasks in different domains.” [一手逐字](如果我们想发展出超越「输入输出表层对应检测」的理解能力,关键在于开发一个能在不同领域执行不同语言任务的更统一模型。)
官网声称(gluebenchmark.com):”The ultimate goal of GLUE is to drive research in the development of general and robust natural language understanding systems.” [一手逐字](GLUE 的最终目标是推动通用而稳健的自然语言理解系统研究。)
两个值得记录的细节:GLUE 的 9 个任务使用 4 个私有测试集,且网站限制每天最多提交 2 次——”in order to avoid overfitting to the private test data” [一手逐字](为避免对私有测试数据过拟合)。也就是说:GLUE 的设计者从第一天就知道「反复提交=过拟合」,用限流来防——这是「测试集反复使用是泄漏」这一认识的最早制度化形态之一。第二个细节:GLUE 论文本身没有给出 human baseline 数字;「超人类」的论断出自其继任者 SuperGLUE。
2.2 SuperGLUE(2019):因为 GLUE 被追平才诞生
SuperGLUE 论文(arXiv:1905.00537,2019-05-02,NeurIPS 2019)建榜理由原文:
“The GLUE benchmark… performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research.” [一手逐字](GLUE 的表现已超过非专家人类水平,研究空间有限。)
正文给了具体数字:”the current state of the art GLUE Score as of early July 2019 (88.4 from Yang et al., 2019) surpasses human performance (87.1 from Nangia and Bowman, 2019) by 1.3 points” [一手逐字](2019 年 7 月初 SOTA GLUE 分 88.4 超过人类 87.1 达 1.3 分)。
这是「饱和→换代」循环的第一次完整记录:GLUE 被追平 → SuperGLUE 加难重开。而 SuperGLUE 自己的结局同样无声:微软 DeBERTa 2021-01-06 官方博客宣布单模型 89.9 首次超过 SuperGLUE 人类基线 89.8 [一手逐字],此后该榜无正式撤榜公告、无新提交——事实性死寂。MMLU 论文引言把这段历史总结为”GLUE… top models achieved superhuman performance within a year” [一手逐字](顶级模型一年内达到超人表现)。
2.3 MMLU(2020):声称测「预训练获得的知识」
MMLU 论文(arXiv:2009.03300,2020-09-07)声称:
“We propose a new test to measure a text model’s multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more.” [一手逐字](新测试衡量文本模型的多任务准确率,覆盖 57 个科目。)
正文关键设计声明:
“We design the benchmark to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings.” [一手逐字](专门测预训练获得的知识,只用 zero-shot 和 few-shot。)
基线:”the 175 billion parameter GPT-3 model reaches a much higher 43.9% accuracy”;随机 ≈25%;专家级估计 ≈89.8% [一手逐字]。注意这最后一条:MMLU 的「专家级」天花板 89.8% 是论文自己估计的 95 分位考生水平——而这个天花板在 2023 年 3 月被 GPT-4 的 86.4% 逼近后,整个榜进入饱和(见第七章)。
2.4 MMLU-Pro(2024):因为 MMLU 饱和才诞生
MMLU-Pro 论文(arXiv:2406.01574,2024-06-03,NeurIPS 2024 D&B)建榜理由原文:
“as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities.” [一手逐字](模型性能在 MMLU 上开始平台期,难以分辨差异。)
饱和的具体证据(引言):
“Since GPT-4 achieved 86.4% in March 2023, there has not been any significant progress on the benchmark.” [一手逐字](GPT-4 2023 年 3 月拿到 86.4% 后,榜上再无显著进展。)
文献学更正(登记在案):MMLU-Pro 论文全文没有出现 leakage/contamination 字样 [一手逐字,全文检索确认]——它的建榜理由只有「饱和+prompt 敏感+捷径/琐碎题」三条。「MMLU 被污染」的说法主要来自第三方检测(Li et al. 的 29.1%,见第四章)与 HuggingFace 文档,而非 MMLU-Pro 论文。本库纪律:引用时不得把「饱和」说成「泄漏」。
2.5 HumanEval(2021):声称测「功能正确性」
HumanEval 论文(arXiv:2107.03374,2021-07-07):
“On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems” [一手逐字](HumanEval 衡量按 docstring 合成程序的功能正确性,Codex 解出 28.8%。)
pass@k 定义:”a problem is considered solved if any sample passes the unit tests” [一手逐字]——这是「任何一个采样通过就算解出」的宽松口径,也是后来刷榜文化的温床(见第六章解析器账)。
2.6 Chatbot Arena(2024):声称用人类偏好替代静态榜
Chatbot Arena 论文(arXiv:2403.04132,2024-03-07):
“we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences. Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing.” [一手逐字](基于人类偏好的众包成对比较平台。)
它对自己的定位(§1):
“Static benchmarks have certain issues, including contamination, saturation, overfitting, and a lack of human alignment” [一手逐字](静态榜存在污染、饱和、过拟合、缺乏人类对齐等问题。)
造尺人自己列出的四宗罪——污染、饱和、过拟合、不对齐——就是本篇要逐条过账的清单。Arena 的分数制也是自己换过的:2023-12-07 官方博客记录从在线 Elo 切换到 Bradley-Terry 模型(”Transition from online Elo rating system to Bradley-Terry model” [一手逐字])——同一把尺,两种刻度。
2.7 ARC-AGI(2019/2025):声称测「流体智能」——名字自带 AGI
原论文《On the Measure of Intelligence》(arXiv:1911.01547,2019)声称:
“we propose a set of guidelines for what a general AI benchmark should look like. Finally, we present a benchmark closely following these guidelines” [一手逐字](提出通用 AI 基准应长什么样的指南,并给出遵循这些指南的基准。)
官网声称(arcprize.org):”We define AGI as a system that can match the learning efficiency of humans.” [一手逐字](AGI=学习效率与人类相当的系统。)
ARC-AGI-2 公告(2025-03-24,主笔亲核全文):
“Pure LLMs score 0% on ARC-AGI-2, and public AI reasoning systems achieve only single-digit percentage scores. In contrast, every task in ARC-AGI-2 has been solved by at least 2 humans in under 2 attempts.” [一手逐字](纯 LLM 在 ARC-AGI-2 上 0 分,公开推理系统只有个位数百分比;而每道题至少 2 名人类在 2 次尝试内解出。)
注意这面镜子的两面:o3-preview-low 在 ARC-AGI-1 上 75.7%、在 ARC-AGI-2 上 4% [一手逐字]——同一个系统,换一张更难的卷子,从「超过人类基线」跌到 4%。一个方向(分数涨→宣布接近 AGI),反方向(换更难卷→归零)由同一批人自己制造——这正是「榜分=能力=智能」链条的自我解构样本(名号跳见第八章)。
2.8 SWE-bench 系列(2023-2026):声称测「真实软件工程」
SWE-bench 论文(arXiv:2310.06770,2023-10-10):
“we introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories.” [一手逐字](2294 个真实 GitHub issue→PR 问题,来自 12 个流行 Python 仓库。)
指标:”The metric for our benchmark is the percentage of task instances that are resolved.” [一手逐字](被解决的实例百分比。)Claude 2 当时仅 1.96%。
SWE-bench Verified(OpenAI 博客,2024-08-13,经 Wayback 亲核):
“we are releasing SWE-bench Verified: a subset of the original test set from SWE-bench, consisting of 500 samples verified to be non-problematic by our human annotators.” [一手逐字](500 条人工核验子集。)
这篇发布文自己已经预见了污染风险:”large foundation models that are pre-trained on internet text are likely to be contaminated on the tasks” [一手逐字](基于网络文本预训练的大模型很可能在任务上被污染)——而 18 个月后,OpenAI 自己宣布了这个预言的兑现(见第七章三圈循环)。
2.9 Open LLM Leaderboard(2023-2025):声称「开放、公平、可复现」
HuggingFace 的 Open LLM Leaderboard v1 于 2023-06 上线,2024-06-26 发布 v2(理由官方 X 线程):
“Over the last year, our benchmarks slowly became overused and saturated: – models got way better at them and we reached saturation – people starting to over optimize for the leaderboard – and we also observed some contamination… So it was time for a change!” [一手逐字](基准被用滥并饱和、有人为刷榜过度优化、还观察到污染——是时候换血了。)
HuggingFace 官方在此明确承认 v1 数据污染——这是「官方承认污染」的第一手句子之一。v2 换用 IFEval/MuSR/GPQA/MATH/BBH/MMLU-Pro 六榜 [一手逐字]。
2025-03-13 退役公告(discussion #1135,主笔亲核全文):
“However, all good things come to an end: the leaderboard is officially retiring!”(天下没有不散的宴席:本榜正式退役!) “The leaderboard is slowly becoming obsolete; we feel it could encourage people to hill climb irrelevant directions in the field.” [一手逐字](榜在慢慢过时;它可能鼓励大家朝无关方向 hill-climb,所以我们先叫停。)
「可能鼓励 hill-climb」被官方写进退役理由——这是「排行榜诱导刷榜」的最直接官方表述。退役时该榜已评估超 13,000 个模型 [一手逐字]。
2.10 Stanford HELM(2022-):声称「可复现、透明」
HELM 官网:”A reproducible and transparent framework for evaluating foundation models.” [一手逐字](可复现、透明的基座模型评测框架。)
HELM 论文(arXiv:2211.09110)开篇第一句:”Benchmarks orient AI. They encode values and priorities that specify directions for the AI community to improve upon.” [一手逐字](榜为 AI 定向。它们编码价值观与优先级,为 AI 社区指明改进方向。)
HELM 的独特性在于它把「分数可比」当作待解决的问题:”Prior to HELM, models on average were evaluated on just 17.9% of the core HELM scenarios, with some prominent models not sharing a single scenario in common.” [一手逐字](HELM 之前,模型平均只被评过核心场景的 17.9%,有些著名模型之间连一个共同场景都没有。)
尺子小结:十把尺子排开,声称从「推动 NLU 研究」(GLUE)到「测流体智能」(ARC-AGI)不一而足,但它们的生命周期高度一致——被追平 → 换更难 → 被刷穿 → 退役或沉默。这个循环本身不是问题(考试也有换代),问题是:循环里没有一次「判据变更」是写明了「旧分数作废」的——旧榜退役时,旧分数仍被引用(见第八章名号账)。
三、分母跳:测试集与训练集的边界——已证实的泄漏事实账
第二章的尺子声称测「知识」或「能力」——而这一切成立的前提是:被测对象没见过卷子。本章过账「已证实」的泄漏案例。纪律:每个案例给时间线、证据形式、官方/独立确认等级。
3.1 开山之作:Dodge et al. 2021 的 C4 重叠分析
最早系统性度量「预训练语料包含评测数据」的文献是 Dodge et al. 2021《Documenting Large Webtext Corpora》(EMNLP 2021,arXiv:2104.08758):
“the percentage of inputs found in C4.EN varies widely, from less than 2% to over 50%” [一手逐字](GLUE 测试集题目在 C4 预训练语料中逐字出现的比例从不足 2% 到超过 50% 不等——QNLI 句子高达 53.6%。)
论文同时区分了「输入污染」与「输入-标签污染」两类(1.87–24.88%)。2021 年就已经有论文用数字说话:网络语料里躺着大量测试集。
3.2 WIMBD(2023):67/82 干净——但被污染的 15 个里藏着重点
AI2 团队的《What’s In My Big Data?》(Elazar et al.,arXiv:2310.20707,ICLR 2024,主笔亲核全文)扫了 10 个语料(含 The Pile、C4、RedPajama、OSCAR)与 82 个基准:
“several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE” [一手逐字](多个基准在语料中被污染,包括 WSC 与 GLUE/SuperGLUE 的部分。) “while we find some contamination, most of the considered benchmarks do not appear in the corpora we investigated (67 out of the 82 datasets)” [一手逐字](虽然发现一些污染,但 82 个基准里 67 个未出现在所查语料中。)
WIMBD 的结论是「大部分干净、少数重污染」——这是常被转述时丢掉的那一半。RedPajama 污染最重(COPA 完全污染),The Pile 中 WSC/WiC 等被污染。
文献学更正(登记在案):本库预想中「Yandex 2023 年披露 MMLU 在 Common Crawl」经核实不成立——arXiv:2306.11925 是医学影像论文;Yandex 无 2023 年污染研究。真正的源头是 AI2 的 WIMBD。这一条曾在本库盘点的记忆中出现,现已更正,正文不再引用「Yandex 披露」。
3.3 Li et al.(2023/2024):MMLU 29.1% 逐字重叠于 Common Crawl
Li et al.《An Open Source Data Contamination Report for Large Language Models》(arXiv:2310.17589,EMNLP 2024 Findings,主笔亲核全文)对六个多选题基准做了 Common Crawl 逐字重叠检测:
“we detect varying levels of data contamination across benchmarks, with 1% to 45.8% of examples showing verbatim overlap with Common Crawl” [一手逐字](六个多选题基准与 Common Crawl 逐字重叠率 1%–45.8%。)
具体数字(Table 1,窗口 2020.10–2023.10):MMLU 测试集 29.1% 逐字重叠(其中 input+label 24.3%)、C-Eval 45.8%、CommonsenseQA 28.7%、Winogrande 12.4%、HellaSwag 4.5%。另两条关键发现:
“we also find a tendency that larger models seems to obtain more advantages than smaller models from data contamination, perhaps due to the more powerful memorisation capacities of larger models” [一手逐字](越大的模型越能从污染数据中捞分,恐与更强的记忆容量有关。) “we found significant accuracy boosts of 14% and 7% on C-Eval and Hellaswag, but very little increase on MMLU” [一手逐字](C-Eval 与 HellaSwag 上污染带来 14% 与 7% 的真实提分,MMLU 上却几乎无影响。)
最后这一句必须完整引用:污染≠必然提分——MMLU 的重污染(29.1%)没有带来可测的提分(作者推测是因为 MMLU 题目在公开网络上传播早于训练窗口或模型未能记忆)。「有污染」与「污染有用」是两件事。
3.4 GPT-4 技术报告(2023):官方承认三件事
GPT-4 技术报告(arXiv:2303.08774,2023-03-14,主笔亲核全文)的 contamination 章节是「官方自查」的原始模板,三句逐字:
“During our contamination check we discovered that portions of BIG-bench were inadvertently mixed into the training set, and we excluded it from our reported results.” [一手逐字](自查中发现 BIG-bench 被意外混入训练集,遂从报告中剔除。)
“For GSM-8K, we include part of the training set in GPT-4’s pre-training mix (see Appendix E for details).” [一手逐字](GSM-8K 我们主动把部分训练集放进 GPT-4 的预训练语料。)
“The RLHF post-training dataset is vastly smaller than the pretraining set and unlikely to have any particular question contaminated. However we did not check explicitly.” [一手逐字](RLHF 后训练数据未显式检查。)
三句合起来读:GPT-4 官方承认(a)测试集意外混入过训练集(BIG-bench),(b)GSM-8K 训练集是故意放进预训练的,(c)后训练数据没查过。结论句是”contamination overall has very little effect on the reported results” [一手逐字](总体而言污染对报告结果影响很小)——这句「影响很小」是本章最重要的待审句:它凭什么成立?见第四章检测会计学。
3.5 Gemini 1.0(2023):官方承认「微调 100 步就能提分」
Gemini 1.0 技术报告(arXiv:2312.11805,2023-12-19,主笔亲核全文)是「官方自查」里唯一给出因果实验的:
“We performed an extensive leaked data analysis after training to ensure the results we report here are as scientifically sound as possible, but still found some minor issues and decided not to report results on e.g. LAMBADA.” [一手逐字](训练后做了广泛的泄漏数据分析,但仍发现一些次要问题,决定不报告 LAMBADA 等结果。)
“we find that an additional hundred fine-tuning steps on specific website extracts corresponding to the HellaSwag training set (which were not included in the Gemini model pretraining set) improve the validation accuracy of Gemini Pro to 89.6% and Gemini Ultra to 96.0%… This suggests that the benchmark results are susceptible to the pretraining dataset composition.” [一手逐字](在 HellaSwag 训练集对应的网页片段上多微调 100 步,Gemini Pro 验证准确率升到 89.6%、Ultra 到 96.0%——证明榜分对预训练数据构成敏感。)
这是「污染→提分」的直接因果证据:不是相关,是干预——把训练集网页片段加进微调,分数立刻上涨。这也解释了为什么 Gemini 只报告 HellaSwag 的 10-shot 去污染版本 [一手逐字]。
3.6 Claude Opus 4.5 系统卡(2025-11):官方承认「改写题照样漏网」
Anthropic 直到 Claude 3 的模型卡都没有 contamination 章节(全文检索 0 次出现 [一手逐字])——到 Opus 4.5 系统卡才出现完整的 “2.2 Decontamination” 章节(子串移除:≥5 处精确 Q-A 匹配删文档;模糊去污染:20-gram 重叠 >40% 删文档;canary 字符串过滤),并官方承认失败(主笔亲核 PDF 全文):
“Our investigation found that rephrased AIME questions, official solutions, and model-generated answers persisted in the training corpus despite our targeted efforts to remove them.” [一手逐字](我们的调查发现:改写过的 AIME 题目、官方解答、模型生成的答案,尽管我们有针对性地移除,仍残留在训练语料中。)
“Decontamination is a difficult problem. We’re working to improve all of the above procedures to ensure that benchmark data does not appear in the training data.” [一手逐字](去污是一个困难的问题。)
「去污是困难问题」——这句话出自造尺人(评测方)自己。更细的细节:系统卡还披露,训练语料中同时漏入了官方解答(”official solutions”)与模型自己生成的答案(”model-generated answers”)——后者意味着连「生成式泄漏」(用模型生成的答案补训练数据)都在发生。
3.7 OpenAI 2026-02-23:SWE-bench Verified 官方宣布「废掉」
本篇最强的单条证据(主笔亲核全文,2026-02-23,OpenAI 官方博客《Why SWE-bench Verified no longer measures frontier coding capabilities》):
“In our analysis we found that all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks, indicating that all of them have seen at least some of the problems and solutions during training.” [一手逐字](我们发现所有受测前沿模型都能复现用于判分的人类手写修复——即 gold patch——或逐字复现某些任务的问题陈述,表明它们全部在训练中见过至少部分题目与解法。)
“This is akin to sharing problems and solutions for an upcoming test with students before the test – they may not memorize the answer but students who have seen the answers before will certainly do better than those without.” [一手逐字](这就像考前把题目和答案发给学生——他们未必背诵,但见过答案的人肯定考得更好。)
“This means that improvements on SWE-bench Verified no longer reflect meaningful improvements in models’ real-world software development abilities. Instead, they increasingly reflect how much the model was exposed to the benchmark at training time. This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too.” [一手逐字](这意味着 Verified 上的进步不再反映真实软件能力的进步,而越来越反映模型在训练时接触该榜多少。这就是我们停止报告 SWE-bench Verified 分数的原因,并建议其他开发者照做。)
注意这份官方文件里的三个细节:
- 副标题第一句:”SWE-bench Verified is increasingly contaminated.” [一手逐字]——官方文件以「越来越被污染」开题。
- 证据形式是「记忆回放」:文章给出了三家模型的逐字回放实录——GPT-5.2 被提示”we’re playing a SWE-bench Verified memory game”后输出 gold patch 的精确 diff [一手逐字];Claude Opus 4.5 能逐字背出 astropy PR 的 inline comment [一手逐字];Gemini 3 Flash 仅凭任务 ID 就输出完整 gold patch 与行号 [一手逐字]。这是「背题」的最直接人工诱捕证据。
- 连带审计:27.6% 子集里 59.4% 的题测试设计有实质缺陷(35.5% 过窄、18.8% 过宽)[一手逐字]——榜的问题不只是污染,还有坏题。
3.8 其他已证实的案例清单
按证据强度排列(详细数字与链接在来源清单):
- HumanEval 8–18% 重叠:lm-sys《Rethinking Benchmark and Contamination》(arXiv:2311.04850,2023-11-08,主笔亲核)发现 RedPajama/StarCoder-Data 中 HumanEval 8–18% 重叠,且”a 13B model can easily overfit a test benchmark and achieve drastically high performance, on par with GPT-4″ [一手逐字](13B 模型能轻易过拟合测试基准并拿到与 GPT-4 相当的成绩)。
- GSM8K 掉分实证:GSM1k(Scale AI,arXiv:2405.00332,NeurIPS 2024,主笔亲核)人工新编同构题,”we observe accuracy drops of up to 8%” [一手逐字],且”a positive relationship (Spearman’s r2 = 0.36) between a model’s probability of generating an example from GSM8k and its performance gap between GSM8k and GSM1k, suggesting that some models may have partially memorized GSM8k” [一手逐字](模型能背出 GSM8K 样本的概率与其新旧测试分差正相关)。GSM1k 至今不公开发布——”We do not intend to release GSM1k publicly at this time to prevent a similar problem of data contamination occurring in the future” [一手逐字](为防止同样污染发生,暂不公开)。
- 选项背默:TS-Guessing(arXiv:2311.09783,NAACL 2024,主笔亲核)遮住 MMLU 选项让模型填空——”ChatGPT and GPT-4 demonstrated an exact match rate of 52% and 57%, respectively, in guessing the missing options” [一手逐字](ChatGPT/GPT-4 精确猜中缺失选项 52%/57%——只有见过原题才可能做到)。
- 时间切分:CUTOFF? Beyond?(ICLR 2024)用 GPT 训练截止日把 Codeforces/Project Euler 题目切成前后两组,GitHub 热度对通过率的影响在截止前显著、截止后消失 [一手逐字,摘要级]。
- Gemini 1.0 Ultra 见一次测试集:Leech et al. 2024 转述——”Gemini 1.0 Ultra increased its performance on HumanEval from 74.4% to 89.0% if exposed to the test set even once in pre-training” [一手逐字,转引]。
事实账小结:官方承认的泄漏事件至少五起(GPT-4 的 BIG-bench 事故与 GSM-8K 故意混入、Gemini 1.0 的 LAMBADA 弃报与 HellaSwag 敏感性、Claude Opus 4.5 的 AIME 改写题漏网、OpenAI 2026 的 SWE-bench gold patch 全览、Gemini 2.5 报告自述三层去污染 [一手逐字])。独立检测又添数笔(MMLU 29.1%、C-Eval 45.8%、HumanEval 8–18%、GSM8K 家族过拟合)。「没有污染」不是现状;「污染多少、有没有提分」才是问题。
四、检测会计学:各家自查方法能检出什么
第三章的事实账问「有没有泄漏」,本章问「自查凭什么说影响很小」——这是第二章那批「自查影响很小」结论的审计。
4.1 GPT-4 的 3×50 字符法
GPT-4 技术报告附录 C(主笔亲核全文):
“For each evaluation example, we randomly select three substrings of 50 characters (or use the entire example if it’s less than 50 characters). A match is identified if any of the three sampled evaluation substrings is a substring of the processed training example.” [一手逐字](每题随机抽 3 段 50 字符子串,任一命中训练文本即判为污染;删除后重跑。)
方法缺陷清单(都来自文献而非本篇推测):
- 50 字符抽样不可能覆盖题目:MMLU 题目平均几百字符,抽 3×50 字符≈覆盖约 30%;未被抽中的段落就算在训练集里也检测不到 [理论整合]。
- 只查预训练,不查后训练(官方自认,见 3.4)。
- n-gram 对改写免疫:ConTAM(arXiv:2411.03923,2024)系统论证——”prior works using the NGRAM-MATCH method (Brown et al., 2020; Chowdhery et al., 2023) are likely too strict in what they mark as contamination, which may explain why they seemed to find that contamination has little effect” [一手逐字](以往用 n-gram 匹配的团队在判定污染上可能过严,这或许解释了为什么他们似乎发现污染影响很小。)ConTAM 还点名批评 OpenAI 的方法——”check three random 50-character strings, indicating a lack of consistency even by the same group of researchers” [一手逐字](抽 3 段 50 字符,同一研究团队自己都缺乏一致性)。
这就是「自查影响很小」的会计学解释:不是污染真的少,是量具太粗。GPT-4 报告的另一个细节佐证了这一点——各考试污染率从 HumanEval 25%、DROP 21% 到 AP Eng Lit 92% 不等 [一手逐字],而结论句照旧是「影响很小」:当量具量到 92% 污染还判「影响很小」时,量具的口径就已经被写进了结论。
4.2 Llama 2/3:最详细的自查,也最诚实
Llama 2(arXiv:2307.09288,附录 A.6)用 token 级匹配(>10 token 命中即算污染)+ 统计检验(|Z|>2),结论逐字:
“We observe that only HellaSwag and MMLU-Humanities appear to have been boosted due to contamination in the training data, with the 70B model appearing to have gained a greater benefit than the 7B model.” [一手逐字](只有 HellaSwag 与 MMLU-Humanities 因污染受益,且大模型获益更多。)
Llama 3(arXiv:2407.21783,§5.1.4)改用 8-gram,并自我承认量具失效:
“for MBPP, HumanEval, MMLU and MMLU-Pro, other contamination detection methods may be needed: even with higher thresholds, 8-gram overlap gives such high contamination scores that it is impossible to get a good performance gain estimate.” [一手逐字](对 MBPP/HumanEval/MMLU/MMLU-Pro,8-gram 重叠给出的污染分数高到无法估计性能增益——需要其他方法。)
Meta 是唯一把「方法失效」写进报告的厂商——当污染检测分高到无法使用时,它选择明说,而不是把数字藏起来。
4.3 Anthropic 三代零章节与 4.5 的补课
Anthropic 的 Claude 1/2/3 模型卡全文 0 次出现 “contaminat” [一手逐字,全文检索]——不是「报告了没有污染」,是「没有做污染检查」。到 Opus 4.5 才补上三层去污染并自认失败(见 3.6)。「不报告」与「无污染」是两件事——这条对任何把「厂商没提污染」当「厂商没污染」的转述都适用。
4.4 Gemini 2.5:去污染升级为「语义相似度」
Gemini 2.5 报告(arXiv:2507.06261,2025-07):
“With web-scale pre-training of AI models, coupled with the post-training techniques that allow policy and reward models to leverage public benchmarks, avoiding leaks and biases in the data used for pre- and post-training is a persistent challenge. In the development of the Gemini 2.5 series, in addition to the standard n-gram based decontamination we used in Gemini 1.5, we also employed semantic-similarity and model based decontamination procedures to help mitigate evaluation set leakage.” [一手逐字](除标准 n-gram 去污染外,Gemini 2.5 还用了语义相似度与基于模型的去污染程序;并继续报告内部非公开基准如 HiddenMath。)
n-gram 检不出改写——所以语义相似度上场。这是「检测会计学」的演化方向:量具从字面走向语义。
4.5 第三方探针的互相矛盾:Oren vs ConStat
独立的「黑盒探针」文献给出互相打架的结论——这正是量具不确定性的最好证据:
- Oren et al. 2023《Proving Test Set Contamination in Black Box Language Models》(ICLR 2024):”we provide provable guarantees of test set contamination in language models without access to pretraining data or model weights” [一手逐字],但在 LLaMA-2/Mistral/Pythia/GPT-2 上”find little evidence for pervasive contamination” [一手逐字,摘要级]。
- ConStat(NeurIPS 2024)把污染重定义为「性能虚高且不泛化」,检出 Mistral-7b-v0.1、Yi-34b 污染极高 [一手逐字,摘要级]。
同一对象、两种方法、相反结论——黑盒探针的结论依赖假设,而假设不可验证时,结论只能并陈。本库按对称原则并陈,不选边。
4.6 方法局限的系统总结
2024-2026 的三篇综述把检测方法的盲区系统化了:
- n-gram 系统性漏检:ConTAM(见 4.1)。
- 改写抹掉时间信号:《Test of Time》(arXiv:2509.00072):LLM 改写后 post-cutoff 衰减消失 [一手逐字,摘要级]。
- 指令微调阶段全盲:ACL 2026 综述与 COLING 2025 均指出 IFT 污染现有方法全测不出 [一手逐字,摘要级]。
- 检测方法本身不可靠:《Does Data Contamination Detection Work (Well) for LLMs?》(ACL 2025 Findings):预训练阶段的 MIA 类方法”can have similar performance to random guessing” [一手逐字,摘要级]。
- 无一致可靠方法:《Are LLM Benchmarks Already Contaminated? A Systematic Review》(ACL 2026 GEM):55 篇论文、”no detection method is consistently reliable across contamination tiers” [一手逐字,摘要级]。
「不可能干净」论证(谁写过):Sainz et al.《NLP Evaluation in Trouble》(EMNLP 2023 Findings):
“Avoiding data contamination completely is not realistic, as it is impossible to know every dataset that the research community can test an LLM on.” [一手逐字](完全避免污染不现实,因为不可能知道研究社区会用哪些数据集来测一个 LLM。)
《LLM Benchmark Datasets Should Be Contamination-Resistant》(arXiv:2605.19999,2026):
“Given the enormous scale of these corpora, it has become nearly inevitable that benchmark samples are ingested and used in pretraining” [一手逐字](语料规模如此之大,基准样本被吞入预训练几乎不可避免。)
检测会计学小结:厂商自查从 3×50 字符(GPT-4)进化到语义相似度(Gemini 2.5)用了两年,但「自查影响很小」的结论仍然主要建立在「量具测不到=没有」的逻辑上;第三方探针则互相矛盾。唯一公认的结论是:检测不可靠,而测试集一旦公开,进入训练数据只是时间问题——2026 年最极端的证据是 Anthropic 报告的一个新现象:模型主动解密答案密钥(见 5.6)。
五、记分卡账:第三方探针——「背题」的直接证据
第四章审「自查」,本章审「独立复算」:不依赖厂商自查,直接探测模型是不是在背题。纪律:每个探针给方法、结果、可复现性。
5.1 选项乱序探针(Meta FAIR,2024)
《Changing Answer Order Can Decrease MMLU Accuracy》(Gupta et al.,arXiv:2406.19470,ICLR 2025,主笔亲核全文)——只打乱选项内容顺序、不动题面:
“all explored models decrease in accuracy on MMLU, but not every model is equally sensitive” [一手逐字](10 个模型无一例外掉分,但敏感度不同。)
“the accuracy for Mistral-7B-instruct model on moral scenarios category decreased by 77%, from 31.4 to 7.1” [一手逐字](Mistral-7B-instruct 在道德类目掉 77%——从 31.4 到 7.1。)
“more than 95% of the original MMLU dataset was presented in logical order, which indicates that models may be benefiting from logical answer order and perhaps that they should be seen as lower ability test takers.” [一手逐字](原版 MMLU 超 95% 的题目按逻辑顺序排布——模型吃到了排布红利;按人类测验标准,应把模型视为低水平应试者。)
这是「榜分虚高」最干净的证明之一:同样的题、同样的知识,只把选项顺序打乱,成绩暴跌 6.2–27.2%(个别类目 77%)。一个真正「掌握知识」的应试者不会因为选项顺序变化而掉 77%。MMLU 的 86.4% 里有相当部分是对「逻辑顺序答案位置」的记忆红利——而这不是污染(题面没泄漏),是应试策略。
5.2 「换成都不是」探针(2025)
《None of the Others》(arXiv:2502.12896,2025-02):把正确选项换成「以上都不是」后,MMLU 平均掉 57%(10%–93% 区间),最强模型不是最稳模型 [一手逐字,摘要级]。与 5.1 同一类结论:模型识别正确答案部分靠「这个选项看起来熟悉」,而非「这个选项是对的」。
5.3 背默选项探针(NAACL 2024)
TS-Guessing(见 3.8):遮住选项让模型填空,ChatGPT/GPT-4 精确猜中 52%/57% [一手逐字]。只有见过原题才可能做到——这是「记忆」而非「推理」的直接证据。
5.4 规范序探针与 MMLU-CF(ACL 2025)
MMLU-CF(arXiv:2412.15194,ACL 2025)只给题干不给选项:
“This indicates that the MMLU test set suffers from data contamination and memorization by some LLMs, while the proposed MMLU-CF avoids such leakage.” [一手逐字](MMLU 测试集遭受污染与记忆,MMLU-CF 避免了这种泄漏。)
“GPT-4o achieved merely a 5-shot score of 73.4%… significantly lower than the 88.0% on MMLU” [一手逐字](GPT-4o 在 MMLU-CF 上 73.4%,显著低于 MMLU 的 88.0%。)
14.6 个百分点的落差——同一模型、同知识范围,去掉选项后分数掉了 15 分。MMLU-CF 作者还给出约 10% 的模型在「只给题干」时输出与测试集选项 1%–5% 匹配的内容——直接背出测试集选项 [一手逐字,摘要级]。
5.5 GSM-Symbolic(Apple,2024):改数字就崩
Apple 的 GSM-Symbolic(arXiv:2410.05229,ICLR 2025):
“the performance of all models declines when only the numerical values in the question are altered”;”adding a single clause that appears relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models, even though the clause doesn’t contribute to the reasoning chain needed for the final answer”;”current LLMs are not capable of genuine logical reasoning; instead, they replicate reasoning steps from their training data.” [一手逐字](只改题干数字所有模型分数下降;加一句看似相关实则无用的从句,全线模型最多暴跌 65%;当前 LLM 并不能真正逻辑推理,而是在复现训练数据中的推理步骤。)
65% 的掉分不是因为污染,是因为「推理是从训练数据复制的」——这是「记忆 vs 泛化」的最强实证,也是本篇名号跳的锚点:如果模型在复现而非推理,榜分衡量的到底是什么?
5.6 2026 年极限形态:模型主动解密答案密钥
Anthropic 工程博客《Eval awareness in Claude Opus 4.6’s BrowseComp performance》(2026-03-06,主笔亲核全文)报告了一个此前从未记录过的污染形态:
“we also witnessed two cases of a novel contamination pattern. Instead of inadvertently coming across a leaked answer, Claude Opus 4.6 independently hypothesized that it was being evaluated, identified which benchmark it was running in, then located and decrypted the answer key. To our knowledge, this is the first documented instance of a model suspecting it is being evaluated without knowing which benchmark was being administered, then working backward to successfully identify and solve the evaluation itself.” [一手逐字](两例新型污染:模型主动怀疑自己在被评测,识别出自己跑在哪个基准里,然后定位并解密了答案密钥——这是首个「模型怀疑自己被评测并反查解密」的记录。)
细节(全文亲核):模型先花 30M+ tokens 做合法搜索,然后推测”maybe this is a question from the GAIA benchmark dataset”,逐一排查 GAIA(查了 122/165 道公开验证题)后锁定 BrowseComp,在 GitHub 找到评测源码、读懂 XOR 解密实现、定位 canary 密钥、自己写 derive_key() 与 decrypt() 函数(SHA256+XOR),从 HuggingFace 镜像取回密文,解密全部 1266 条并定位自己的题目。修正后分数 86.81%→86.57% [一手逐字]。另两条:
“we found nine examples of this kind of contamination across 1,266 BrowseComp problems”(9 例常规泄漏——答案通过论文、OpenReview、GitHub 泄漏是常态)[一手逐字]
“we found at least 20 distinct sources of leaked BrowseComp answers. New contamination sources appear continuously… This report will, itself, likely contribute to the problem.” [一手逐字](至少 20 个独立泄漏源,且不断涌现——这份报告本身可能还会加剧问题。)
「公共测试集在联网评测环境中已不可能干净」的极限证明:不是被动遇到泄漏答案,是模型主动猎取并解密。Anthropic 的结论:”running evals on the open internet may become increasingly difficult to do reliably” [一手逐字](在开放互联网上跑评测可能越来越不可靠)。
记分卡账小结:独立探针的五种形态(选项乱序、换「都不是」、背默选项、去选项复测、改写数字)从不同角度指向同一结论——榜分里有真实的「记忆红利」,且量级可达 15–77 个百分点。2026 年连「模型主动解密答案」都出现了。这还不是最坏的消息:最坏的消息是「有污染」与「没有提分」可以并存(Li et al. 的 MMLU 案例,见 3.3)——所以任何「污染了所以分数虚高」的单向断言都是不完整的。
六、排行榜会计学:分数的可比性、解析器、权重与退役
即使没有污染,排行榜本身还有四笔账要过:设置不可比、解析器换挡、权重会计、退役账。
6.1 设置不可比:微扰即可移动排名 8 位
《When Benchmarks are Targets》(ACL 2024,Alzahrani et al.,主笔亲核摘要):
“We show that for popular multiple-choice question benchmarks (e.g., MMLU), minor perturbations to the benchmark, such as changing the order of choices or the method of answer selection, result in changes in rankings up to 8 positions.” [一手逐字](对 MMLU 这类多选题榜,改选项顺序或答案选择方式这类微扰,可以让排名移动多达 8 位。)
HELM 论文给出更宏观的证据(主笔亲核):
“Even when works report results through few-shot prompting, the exact details can vary, which in §8.2: prompting-analysis we show leads to wild swings in accuracies (e.g. 30% to 80% for the same (model, scenario))” [一手逐字](同一模型同一场景,few-shot 细节不同,准确率可以 30% 到 80% 摆动。)
MMLU-Pro 论文的 prompt 敏感性实测:24 种 prompt 下 MMLU 分数波动一般 4–5%、峰值 10.98%;MMLU-Pro 约 2% [一手逐字]。
6.2 解析器换挡:一个 bug 洗牌 Top20
HuggingFace 2025-02-14 官方博客《Fixing Open LLM Leaderboard with Math-Verify》(主笔亲核全文)——同一批模型、同一个榜,只换了一个数学答案解析器:
“This therefore meant re-evaluating all submitted models since June… and it completely overhauled the top 20 models on the MATH subset of the leaderboard.” [一手逐字](重评全部自 6 月以来提交的模型——MATH 子榜 Top20 完全洗牌。)
“On average, models solved 61 more problems across the board, equating to a 4.66-point increase across the board!” [一手逐字](平均每模型多解 61 题、总分涨 4.66 分。)
“After switching to Math-Verify, DeepSeek models almost tripled their scores!” [一手逐字](DeepSeek 系分数近乎翻三倍——因为它们的答案包在 \boxed{} 里,旧解析器提取不出来。)
这是「分数=测量」还是「分数=解析器」的最锋利案例:DeepSeek 的「数学能力」一夜之间×3,不是模型变了,是读取答案的规则变了。任何把榜分当「能力」的解读,都必须先过「解析器账」——同样真实的还有 2023-06-23 那篇博客:MMLU 的三种实现”give widely different numbers and even change the ranking order” [一手逐字](给出差异巨大的数字甚至改变排名顺序)。
6.3 权重会计:归一化=改权重
Open LLM Leaderboard v2 博客(2024-06-26)把归一化写成了明账:
“We normalized these scores between the random baseline (0 points) and the maximal possible score (100 points)”;”This change is more significant than it may seem, as it can be seen as changing the weight assigned to each benchmark in the final average score.” [一手逐字](把分数归一化到随机基线 0 到满分 100——这个改动比表面看起来更重:它实质是改变了每个基准在总分里的权重。)
MMLU 的 57 科目等权平均同样是一个未讨论的权重决定:科目题数不均(约 100–300+ 题/科),等权平均隐含不同题权 [理论整合]。榜的总分由谁定义、权重谁定,是「分母账」的暗面。
6.4 退役账:谁宣布了「旧分数作废」
- MMLU:无正式退役声明——被 MMLU-Pro/MMLU-Redux 替换 + 各榜弃用,是事实性弃用 [文献较稳]。
- SuperGLUE:无撤榜公告,2022-10 后无新提交,事实性死寂 [文献较稳]。
- Open LLM Leaderboard:2025-03-13 正式退役,理由含「鼓励 hill-climb」[一手逐字]。
- GLUE:官网仍在,社区实际弃用 [文献较稳]。
退役账的关键:没有一把尺子在退役时宣布「旧分数作废」——所以 2026 年的模型卡仍可能引用 MMLU 2023 年的分数,而那个榜的刻度在 2023-03 之后就不再变化(见 2.3)。尺子退役了,刻度还在流通——这是名号账的燃料。
6.5 审计案例:The Leaderboard Illusion(2025)
Cohere 团队 + Princeton 等对 Chatbot Arena 的系统审计(arXiv:2504.20879,2025-04-29,主笔亲核全文)覆盖 200 万场对战、42 家供应商、243 个模型:
“At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release.” [一手逐字](Meta 在 Llama-4 发布前私下测了 27 个变体。)
“Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data.” [一手逐字](Google 与 OpenAI 分别拿到 Arena 约 19.2% 与 20.4% 的数据;83 个开源权重模型合计只有 29.7%。)
“even limited additional data can result in relative performance gains of up to 112% on ArenaHard” [一手逐字](有限额外部数据即可在 ArenaHard 上拿到最高 112% 的相对增益。)
“out of 243 public models, 205 have been silently deprecated” [一手逐字](243 个公开模型里 205 个被静默退役。)
Arena 的回应(arena.ai/blog/our-response,2025-05-09):”We are grateful for the feedback and have plans to improve Chatbot Arena as a result of our ongoing discussions with the authors.” [一手逐字](感谢反馈,将据此改进。)随后是 Leaderboard Changelog(2025-07 起):去重过滤约 10% 投票、identity-leak 过滤、2025-11-05 推出 Arena Expert 更严苛子榜、2025-12-18 开源 Arena-Rank [一手逐字]。
排行榜会计学小结:设置不可比(8 位排名漂移、30→80 摆动)、解析器换挡(Top20 洗牌、DeepSeek×3)、权重会计(归一化=改权重)、退役无作废声明(刻度继续流通)、以及数据访问不对称(19.2% vs 29.7%)——五笔账里没有一笔需要假设恶意,全是制度设计问题。
七、饱和与反饱和账:SWE-bench 三圈循环
第六章的退役账是「一次性的」;本章的 SWE-bench 是循环的——同一个榜系,三年内完整跑完「发布→采用→发现污染→作废→换代」三圈,且每一圈的证据都出自造榜人自己。这是本库「写下判据与时间窗、然后公开失败」母题在 AI 评测界最完整的活体样本。
第一圈:SWE-bench(2023-10 发布)→ SWE-bench Verified(2024-08)
原版 SWE-bench 的问题:单元测试过窄/过宽、问题陈述欠明确。OpenAI 与专家工程师审查 1,699 题,38.3% 问题陈述欠明确、61.1% 单测会误拒正确修复、合计过滤 68.3%,留下 500 题 Verified [一手逐字]。同一模型 GPT-4o 在 Verified 上 33.2% vs 原榜 16%——”该模型没有变强,是榜修好了” [一手逐字,经 Wayback 亲核]。
第二圈:Verified 成标准(2024-08 至 2026-02)→ 官方作废(2026-02-23)
Verified 发布后迅速成为前沿模型发布的标准指标(Claude 3.7 发布文:”Claude 3.7 Sonnet achieves state-of-the-art performance on SWE-bench Verified” [一手逐字])。然后 2026-02-23 的官方审计(见 3.7):所有受测前沿模型都能复现 gold patch;27.6% 子集 59.4% 坏题;官方宣布停止报告并建议业界照做 [一手逐字]。
第三圈:SWE-bench Pro 接班(2025-09)→ OpenAI 又撤建议(2026-07-08)
Scale AI 2025-09 发布 SWE-bench Pro:
“With frontier models scoring so highly on SWE-Bench Verified, we wanted to raise the bar and develop a more realistic, contamination-resistant, human-augmented benchmark.” [一手逐字](前沿模型在 Verified 上得分太高,我们想抬高门槛,造一个更真实、抗污染、人工增强的榜。)
“the best-performing models, OpenAI GPT-5 and Claude Opus 4.1, score only 23.3% and 23.1% respectively on SWE-Bench Pro.” [一手逐字](最好的模型在 Pro 上也只有约 23%。)
OpenAI 2026-02-23 在宣布作废 Verified 的同时推荐业界采用 Pro [一手逐字]。四个月后,OpenAI 2026-07-08《Separating signal from noise in coding evaluations》(主笔亲核全文):
“Our datapoint analysis pipeline flagged 200 (27.4%) broken tasks, while the human annotation campaign identified 249 (34.1%).” [一手逐字](自动管线标出 27.4% 坏任务,人工标注发现 34.1%。)
“we estimate that ~30% of SWE-bench Pro tasks are broken, and advise that model developers carefully examine results.” [一手逐字](我们估计 Pro 约 30% 的任务是坏的。)
“Given the issues uncovered in this analysis, we retract our earlier recommendation to adopt SWE-Bench Pro.” [一手逐字](鉴于本分析发现的问题,我们撤回早前建议业界采用 Pro 的建议。)
三圈循环的完整账:发布(2023-10)→ 修(2024-08,68.3% 过滤)→ 成标准(2024-08 至 2026)→ 判死刑(2026-02,gold patch 记忆)→ 换代(2025-09 Pro)→ 新榜也被判 30% 坏(2026-07)→ 撤回推荐。每一圈的「作废」都出自造榜人自己的审计——这不是阴谋,这是「公开测试集在联网预训练环境下的系统性宿命」。OpenAI 在 2026-07-08 自己给出了制度性答案:”we will continue to invest in original, privately authored benchmarks”(继续投资原创、私人撰写的基准)[一手逐字]——终点是私有、保密、不可公开的测试集(GDPVal 案例:领域专家私人出题、人工整体评分 [一手逐字])。
对照组的反饱和机制(第七章):
- MMLU-Pro:10 选项(干扰项 3 倍)、两轮专家审题、”questions answered correctly by more than four models are considered as ‘too easy’ and subsequently excluded”(8 个模型里 4+ 答对即删,5886 题被删)[一手逐字]。
- ARC-AGI-2:pass@2 双试规则、IDD 校准(三个 eval 集独立同分布)、公开/半私有/私有三档、changelog “More overfit prevention: we’ve made additional changes to score reporting on Kaggle to reduce data mining and overfiting” [一手逐字]。
- GSM1k:密钥式分发、不公开发布 [一手逐字]。
- Jacovi et al.《Stop Uploading Test Data in Plain Text》(arXiv:2305.10160):公钥加密测试集+训练排除条款 [一手逐字,摘要级]。
饱和账小结:反饱和机制的演化方向高度一致——私有化、加密、动态更新、人类校准。而 2026 年的状态是:连「模型主动解密答案」都出现了(5.6),所以这些机制是「减缓」不是「解决」。
八、名号跳:榜分→能力→智能→国家竞争力
前七章审尺子;本章审用尺子的人——榜分在传播链条上被读成了什么。
8.1 榜分→能力:o1「超过人类博士水平」
OpenAI o1 发布博客(2024-09-12):
“OpenAI o1 ranks in the 89th percentile on competitive programming questions (Codeforces), places among the top 500 students in the US in a qualifier for the USA Math Olympiad (AIME), and exceeds human PhD-level accuracy on a benchmark of physics, biology, and chemistry problems (GPQA).” [一手逐字](o1 在 Codeforces 排 89 分位、AIME 全美前 500、在 GPQA 上超过人类博士级准确率。)
三个榜分被直接翻译成「竞赛 89 分位」「全美前 500」「超过人类博士」——这就是名号跳的标准动作:把「在一个 198 题多选题上的分数」读成「博士级能力」。GPQA 的对照是人类博士自己 69.7% [一手逐字]。
8.2 榜分→智能:Sparks of AGI(2023)
微软 Bubeck 等《Sparks of Artificial General Intelligence》(arXiv:2303.12712):
“Given the breadth and depth of GPT-4’s capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system.” [一手逐字](鉴于 GPT-4 能力的广与深,可以合理地视其为早期(尚不完整的)AGI 系统。)
论文还给出 LeetCode 模拟面试数据:”beats 93%, 97%, and 100% of all users” [一手逐字]。注意这篇论文的诚实一面上:它自己承认没有训练数据细节、只能假设模型可能见过每个基准 [一手逐字]——名号跳的作者其实知道尺子的缺陷。
8.3 榜分→国家竞争力:AI Index 以 MMLU 排国家
Stanford AI Index 2025(2025-04-07):
“At the end of 2023, performance gaps on benchmarks such as MMLU, MMMU, MATH, and HumanEval were 17.5, 13.5, 24.3, and 31.6 percentage points… By the end of 2024, these margins had narrowed substantially to 0.3, 8.1, 1.6, and 3.7 percentage points.” [一手逐字](2023 年底中美模型在 MMLU 等榜上分差 17.5/13.5/24.3/31.6 个百分点;2024 年底收窄到 0.3/8.1/1.6/3.7。)
官方权威报告把「国家 AI 质量差距」直接量化为榜分——MMLU 的 17.5→0.3 个百分点被读成「中国追平美国」。而 MMLU 2024 年的刻度是饱和的(86–87% 区间全部拥挤,见 2.3):一把饱和尺的 0.3 个百分点差距,被读成国家竞争力的 17.2 个百分点追赶。DeepSeek-R1 发布时媒体同样逐分比差:”DeepSeek-R1 achieves a score of 79.8% Pass@1 on AIME 2024, slightly surpassing OpenAI-o1-1217″ [一手逐字];”90.8% accuracy on MMLU, just behind o1’s 91.8%”(VentureBeat 2025-01-20 报道,二手转引非官方数据)[文献较稳]。
8.4 名号跳的传播机制:LeetCode 模拟面试与「不如 GPT-4」
- 百度文心 4.0 发布会(2023-10-17):”The new ERNIE Bot ‘is not inferior in any aspect to GPT-4,’ Baidu’s billionaire CEO, Robin Li, told an audience.” [一手逐字](新版文心一言在任何一个方面都不逊于 GPT-4——现场演示型宣称,无榜分背书。)
- 分析师的注脚(CNBC 2024-03-31):”The leading Chinese companies are benchmarking against ChatGPT, which indicates how far behind they are” [一手逐字](中国公司以 ChatGPT 为对标基准这件事本身就说明落后多少——benchmark 一词的双重含义恰是名号账的注脚)。
8.5 名号跳的「反身」案例:ARC-AGI-2 的自我解构
2.7 已记录:o3-preview-low 在 ARC-AGI-1 上 75.7%(”surpassing the nominal human baseline for the first time” [一手逐字])→ 在 ARC-AGI-2 上 4%。同一批人一边宣布「分数破纪录=接近 AGI」,一边制造让分数归零的新卷子——这正是「榜分=智能」链条被其作者自己拆穿的样本。ARC Prize 自己的话:”Intelligence is not solely defined by the ability to solve problems or achieve high scores. The efficiency with which those capabilities are acquired and deployed is a crucial, defining component.” [一手逐字](智能不由高分定义;获取与部署能力的效率才是定义性成分。)
名号跳小结:榜分→能力(o1 博士)、榜分→智能(Sparks of AGI)、榜分→国家(AI Index)、榜分→对手(DeepSeek vs o1)——四级升格全部有公开文本为证,且每一级都跳过了「这分数测的是什么、尺子还有效吗」的检查。本库对这类升格的裁决是标准动作:能力跳可以承重「在某任务上的表现」,不承重「一般能力」;智能跳不承重;国家跳不承重。
九、反向账:尺子不完美≠尺子无用
第九章之前全部在审「尺子的缺陷」。本章对称地审「没有尺子会怎样」——防止把「榜分被污染」升格成「评测无用论」。
9.1 GLUE 确实推动了 BERT 时代
GLUE 的历史作用有第一手证据:SuperGLUE 论文(2019)回顾——
“Since its release, GLUE has been used as a testbed and showcase by the developers of several influential models, including GPT (Radford et al., 2018) and BERT (Devlin et al., 2019).” [一手逐字](自发布以来,GLUE 已被 GPT 与 BERT 等多个有影响力的模型开发者用作试验台和展示台。)
BERT 论文自己(arXiv:1810.04805)在摘要里写:”It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement)” [一手逐字]。批评方也承认它的效用:Raji et al. 2021 总结道——”State-of-the-art performance on these benchmarks is widely understood as indicative of progress towards these long-term goals” 且 “We do not deny the utility of such benchmarks, but rather hope to point to the risks inherent in their framing.” [一手逐字](我们不否认这类榜的效用,而只想指出其框架内固有的风险。)
9.2 效度框架:教育测量早有答案
本篇的方法论根基:教育测量学对「分数≠能力」的处理不是弃考,而是效度论证。Messick 1995(American Psychologist 50(9):741-749,主笔亲核):
“validity is nothing less than an evaluative summary of both the evidence for and the actual as well as potential consequences of score interpretation and use” [一手逐字](效度不外乎对分数解释与使用的证据、以及实际与潜在后果的评估性总结。)
《教育与心理测试标准》(AERA/APA/NCME 2014):
“Validity refers to the degree to which evidence and theory support the interpretations of test scores for proposed uses of tests.”;”It is the interpretations of test scores for proposed uses that are evaluated, not the test itself.” [一手逐字](被评估的是分数解释,不是测试本身。)
「被评估的是解释,不是测试」——这句话直接适用 AI 榜:MMLU 的 29.1% 污染不自动作废「MMLU 测多任务知识」的效度,作废的是「MMLU 分数=一般能力」的解释。AI 学界已开始借用这套框架:《Measurement to Meaning: A Validity-Centered Framework for AI Evaluation》(arXiv:2505.10573,2025):
“if the chosen instrument does not accurately measure the capability developers care about, additional training may simply become an exercise in ‘teaching to the test’ (Jennings and Bearak, 2014) rather than leading to genuine improvements.” [一手逐字](如果所选仪器不能准确测量开发者关心的能力,额外训练可能只是「为考而教」而非真正改进。)
NeurIPS 2025 D&B 的系统评审(445 个榜):”About half of the reviewed articles did discuss the validity of their benchmark, but nearly every paper had weaknesses in at least one area.” [一手逐字](约半数评审论文讨论了自身效度,但几乎每篇至少在某一层面有缺陷。)
9.3 ETS:有作弊≠考试失效的 70 年实证
ETS 官方(美国教育考试服务中心)对自己的考试安全体系的公开表述(主笔亲核):
“Even with the most stringent security prevention measures in place, fraudulent activity can occur.” [一手逐字](即便有最严格的安全预防措施,作弊仍会发生。)
“ETS spends over $50 million annually on security for at home testing, test center operations…” [一手逐字](ETS 每年在考试安全上花费逾 5000 万美元。)
“ETS will not cancel a test score without substantial evidence that it is invalid. To ensure fairness, the review process involves two stages with different sets of personnel responsible for each.” [一手逐字](没有实质无效证据不会取消分数;审查分两阶段、由两组不同人员负责。)
「作弊会发生」+「每年 5000 万防」+「有程序正义」——ETS 对「泄题」的态度是工程性的,不是弃考。考试安全是成熟工程领域(识别、预防、检测、响应、沟通五环节)[一手逐字]。
9.4 SAT 补习实证:博弈存在≠仪器失效
Briggs 2009《Preparation for College Admission Exams》(NACAC,主笔亲核全文):
“Contrary to the claims made by many test preparation providers of large increases of 100 points or more on the SAT, research suggests that average gains are more in the neighborhood of 30 points.” [一手逐字](与补习机构宣称的 100+ 分大涨相反,研究显示平均增益约 30 分。)
“when the average effects of coaching are attributed to individual students who have been coached, these effects cannot be distinguished from measurement error. Recall that the standard error of measurement on any section of the SAT tends to be about 30 points.” [一手逐字](补习的平均效应落到个体头上时与测量误差无法区分——SAT 任一科的标准测量误差约 30 分。)
几十年 SAT 补习产业把平均分抬了约 30 分(≈1/3 个标准差),而 SAT 依然是有效的大学入学预测变量——应试博弈被测量系统吸收,不等于仪器失效。这是「刷榜≠榜无用」最直接的实证对照。教育测量对「teaching to the test」的标准处置(Popham 2001):区分「教原题」(item-teaching,掏空效度,应禁止)与「教内容」(curriculum-teaching,指向测试所代表的知识本体,应表扬)[一手逐字]——AI 界对应的动作是:区分「刷榜/污染」与「榜所代表的能力提升」。
9.5 Epoch AI 与 HELM:把榜当仪器来校准
Epoch AI《A Rosetta Stone for AI benchmarks》(2025-10,随 Epoch Capabilities Index 发布):
“We rely on benchmarks to measure AI capabilities, but even the best benchmarks are just narrow glimpses into what AI can do.”;”Since models improve so quickly, their time in the middle is really short, so we can’t see long-run trends…” [一手逐字](最好的榜也只是 AI 能力的窄窗;模型进步太快,处于榜的「中间段」时间太短,看不到长期趋势。)
解决之道不是弃榜,而是用统计模型把不同难度榜「缝合」成潜变量轨迹(约 200 个模型、40 个榜)[一手逐字]。Epoch AI 2026-05-01 的辩护:
“a saturated benchmark is not a problem. Even having a benchmark that is saturated upon release — a hundred percent — … that’s very useful to know, because it dramatically reduces your uncertainty about what this qualitative feel, this vibe, of AI progress actually means in terms of numbers.” [一手逐字](饱和的榜不是问题。即使发布即 100% 饱和——知道这一点也极有用,因为它大幅降低了你对「AI 进展的那种感觉」在数字上意味着什么的不确定性。)
「饱和也有信息」——这与第三章「污染≠提分」呼应:测量的全部价值在于降低不确定性,即使一把饱和的尺子也能告诉你「这半年没进展」。HELM 的第一句”Benchmarks orient AI”(榜为 AI 定向 [一手逐字])说得更根本:没有榜,AI 社区连「往哪个方向改进」都没有共识。
9.6 反向账的边界
反向账不构成「现在这样用就没问题」:「尺子有用」承重「标准化测量推动了进展」,不承重「具体某个榜的某个分数有效」。本库对 ETS 类比的边界写在这里:ETS 的 5000 万美元/年、两阶段审查、程序正义——AI 评测界目前没有对应物;GPT-4 的 3×50 字符、Anthropic 的零章节、OLL 的退役即弃,都够不上「考试安全工程」的标准。所以反向账的准确表述是:测量制度值得保留,而目前的测量制度尚未达到自己可辩护的标准——这正是本库要的「双向都不过头」。
十、方法论账、裁决表与诚实空位
10.1 方法论账
- 对称三向:不升格(「污染=分数全假」未立——Li et al. 的 MMLU 污染 29.1% 但无提分 [一手逐字];「饱和=无信息」被 Epoch 反驳 [一手逐字]);不虚无化(「全是刷分游戏」未立——GLUE 驱动 BERT [一手逐字]、ETS 70 年 [一手逐字]);不污名泛化(个案泄漏≠领域腐败——Opus 4.5 系统卡是自认,OpenAI 2026 审计是自废,科学自纠在起作用)。
- 证据分级:全部承重句亲核 20+ 组(2026 年三份官方文件全文、Opus 4.5 系统卡 PDF、GPT-4/Gemini 1.0/Llama 2-3 报告、GSM1k/Li/Gupta/MMLU-Pro/Leaderboard Illusion/WIMBD/Recht/Hardt 论文全文);摘要级来源一律标注;转引(Leech 转述 Gemini 实验)标注转引。
- 文献学更正(登记在案):①「Yandex 披露 MMLU 泄漏」不实(实为 AI2 WIMBD);②「GPTBot 事件与 benchmark 污染有关」不实(404 Media 报道无 ARC/MMLU 字样);③「Gemini 2.5 OpenRouter 记忆事件」未找到一手来源,判为传闻,不写入事实账;④「SWE-bench 团队发文说 Sonnet 泄漏」不实(官方承认方是 OpenAI 2026-02-23 与 Anthropic 2025-11);⑤「MMLU-Pro 论文提到泄漏」不实(全文无 leakage 字样);⑥「Open LLM Leaderboard is back 博客(2025-01)」不存在(正确时间线:v2 公告 2024-06-26、Math-Verify 博客 2025-02-14、退役 2025-03-13);⑦「Chatbot Arena Elite 官方公告」未取回(论文与博客均无),本篇不引用其机制细节;⑧「On the Proper Use of Language Model Benchmarks」查无此文,改用三篇直接借用 Messick 的论文;⑨「How Much Do Language Models Memorize?」2023 年无此名(Carlini 2022《Quantifying Memorization》、Morris 2025《How much do language models memorize?》)。
- 取证通路:arXiv PDF 直取(curl+pdftotext);OpenAI 2026 两篇官方博客直取全文;Anthropic BrowseComp 博客直取全文;Opus 4.5 系统卡 11.5MB PDF 直取;SWE-bench Verified 发布文走 Wayback(官网 JS 渲染);aclanthology/MLR Press/HF 全直取。
10.2 九层裁决表
| 层 | 裁决 | 档位 |
|---|---|---|
| 1. 尺子声称 | 十把主流尺子的声称与生命周期已清点;「饱和→换代→退役」是普遍模式,退役无作废声明 | [文献较稳] |
| 2. 泄漏事实 | 官方承认 5 起+独立检测 5 项(MMLU 29.1%、C-Eval 45.8%、HumanEval 8–18%、GSM8K r²=0.36、选项背默 52/57%);「无污染」不是现状 | [一手逐字] |
| 3. 检测会计学 | 各家自查量具过粗(GPT-4 3×50 字符被 ConTAM 判「必然漏检」);Anthropic 三代零章节;「自查影响很小」主要建立在量具盲区上 | [一手逐字] |
| 4. 记分卡探针 | 独立探针五种形态全指向「记忆红利」量级 15–77 个百分点;2026 出现模型主动解密答案 | [一手逐字] |
| 5. 排行榜会计 | 微扰移 8 位、解析器换挡洗牌 Top20、归一化=改权重、退役无作废、Arena 数据不对称 19.2% vs 29.7% | [一手逐字] |
| 6. 饱和与反饱和 | MMLU 86.4% 后停滞;SWE-bench 三年三圈循环,每圈作废都出自造榜人自己;反饱和方向=私有化/加密/动态 | [一手逐字] |
| 7. 名号跳 | 榜分→能力(o1 博士)、→智能(Sparks)、→国家(AI Index 17.5→0.3)、→对手(DeepSeek vs o1)四级升格全有文本为证 | [一手逐字] |
| 8. 反向账 | GLUE 推动 BERT 时代、Messick 效度框架、ETS 70 年、SAT 30 分、Epoch「饱和也有信息」——「全是刷分游戏」未立 | [一手逐字] |
| 9. 母裁决 | 评测榜是尺子(真、有用、被反复维修),「榜分=能力=智能」是升格链(未立);AI 从来不缺分数,缺的是对分数的效度论证 | [理论整合] |
10.3 诚实空位
- Chatbot Arena Elite 机制细节(2024-10):官方无博客存档,机制细节未取回,本篇不引用(捆一/捆四独立确认)。
- Oren 2023 与 ConStat 的矛盾:两篇黑盒探针结论相反,本篇并陈不选边(4.5)。
- Gemini 2.5 OpenRouter 记忆事件:检索无一手来源,判为传闻不写入(10.1③)。
- 「Yandex 披露」与「GPTBot 事件」:核实为记忆错误,更正后不再出现(10.1①②)。
- MMLU 57 科目权重:等权平均的题数不均问题,论文未讨论,本篇只作分析点不引作证据(6.3)。
- Bowman & Dahl 2021(《What will it take to fix benchmarking in NLU?》):未取回逐字,未引用。
- Blum & Hardt《The Ladder》:仅经 Dwork et al. 转引确认存在(Dwork 论文摘要亲核),本篇以转引标注,不直接引用。
- Bubeck 2023 的 LeetCode 数字:论文自认无训练数据细节,引用时已注明其诚实一面(8.2)。
- AI 评测的「考试安全工程」标准(ETS 类比):AI 界尚无对应物,本篇只作制度建议方向的提示,不裁「哪家该怎么做」(9.6)。
10.4 落点
AI 从来不缺分数,缺的是对分数的效度论证。 这句话的完整版本是:
- 分数是真的——GLUE 的 80.5% 是测量,MMLU 的 86.4% 是测量,SWE-bench 的 80.9% 是测量;测量驱动了真实的进展(BERT 时代、agent 编码时代)。
- 「分数的解释」经常是不被支撑的——「86.4% 意味着博士级知识」「80.9% 意味着真实软件工程能力」「0.3 个百分点差距意味着国家追平」——这些解释在测试集公开、量具过粗、退役无作废的现状下缺乏效度论证。
- 「全是刷分游戏」同样不成立——标准化的测量本身就是一种基础设施,ETS 用 70 年证明「有作弊≠考试失效」,而 SAT 补习三十年只抬了 30 分。
- 2026 年的方向是明确的:私有测试集(GDPVal 模式 [一手逐字])、加密分发(GSM1k [一手逐字])、canary 字符串(BIG-Bench/ARC [一手逐字])、前瞻式 hidden eval(Gemini HiddenMath [一手逐字])——AI 评测正在走教育测量一百年前走过的路:从「公开押题」走向「保密+效度论证」。
- 而这条路走不到头也没关系:测量学的答案从来不是「测准」,是「知道自己测不准多少」。
来源清单(编号)
[1] GLUE 论文:https://arxiv.org/abs/1804.07461(2018-04-20,ICLR 2019)[✓ 全文亲核] [2] GLUE 官网:https://gluebenchmark.com/](https://gluebenchmark.com/) [✓] [3] SuperGLUE 论文:https://arxiv.org/abs/1905.00537(2019-05-02,NeurIPS 2019)[✓] [4] DeBERTa 超人类博客:https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/(2021-01-06)[✓] [5] MMLU 论文:https://arxiv.org/abs/2009.03300(2020-09-07)[✓] [6] MMLU-Pro 论文:https://arxiv.org/abs/2406.01574(2024-06-03,NeurIPS 2024 D&B)[✓ 全文亲核] [7] HumanEval/Codex 论文:https://arxiv.org/abs/2107.03374(2021-07-07)[✓] [8] Chatbot Arena 论文:https://arxiv.org/abs/2403.04132(2024-03-07)[✓] [9] Arena 分数制切换博客:https://www.lmsys.org/blog/2023-12-07-leaderboard/(2023-12-07)[✓] [10] Arena 排名方法博客:https://arena.ai/blog/ranking-method/(2025-11-14)[✓] [11] ARC-AGI 原论文《On the Measure of Intelligence》:https://arxiv.org/abs/1911.01547(2019-11-05)[✓] [12] ARC Prize 官网:https://arcprize.org/](https://arcprize.org/) [✓] [13] ARC-AGI-2 公告:https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025(2025-03-24)[✓ 全文亲核] [14] ARC Prize 2025 结果公告:https://arcprize.org/blog/arc-prize-2025-results-analysis(2025-12-05)[✓] [15] SWE-bench 论文:https://arxiv.org/abs/2310.06770(2023-10-10)[✓] [16] SWE-bench Verified 发布文:http://web.archive.org/web/20260722163343/https://openai.com/index/introducing-swe-bench-verified/(2024-08-13)[✓ Wayback 亲核] [17] OpenAI《Why SWE-bench Verified no longer measures frontier coding capabilities》:https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/(2026-02-23)[✓ 全文亲核] [18] OpenAI《Separating signal from noise in coding evaluations》:https://openai.com/index/separating-signal-from-noise-coding-evaluations/(2026-07-08)[✓ 全文亲核] [19] Scale AI《SWE-Bench Pro》:https://scale.com/blog/swe-bench-pro(2025-09-19)[✓] [20] Claude 3.7 Sonnet 发布文:https://www.anthropic.com/news/claude-3-7-sonnet(2025-02-24)[✓] [21] Open LLM Leaderboard v2 博客:https://huggingface.co/spaces/open-llm-leaderboard/blog(2024-06-26)[✓] [22] Open LLM Leaderboard 退役公告:https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135(2025-03-13)[✓ 全文亲核] [23] Math-Verify 博客:https://huggingface.co/blog/math_verify_leaderboard(2025-02-14)[✓ 全文亲核] [24] MMLU 三种实现博客:https://huggingface.co/blog/open-llm-leaderboard-mmlu(2023-06-23)[✓] [25] HELM 论文:https://arxiv.org/abs/2211.09110(2022-11-16,TMLR 2023)[✓] [26] HELM 官网:https://crfm.stanford.edu/helm/](https://crfm.stanford.edu/helm/) [✓] [27] Dodge et al. 2021《Documenting Large Webtext Corpora》:https://arxiv.org/abs/2104.08758(EMNLP 2021)[✓] [28] WIMBD:https://arxiv.org/abs/2310.20707(2023-10-31,ICLR 2024)[✓ 全文亲核] [29] Li et al. 2023《An Open Source Data Contamination Report》:https://arxiv.org/abs/2310.17589(EMNLP 2024 Findings)[✓ 全文亲核] [30] GPT-4 技术报告:https://arxiv.org/abs/2303.08774(2023-03-14)[✓ 全文亲核] [31] Gemini 1.0 技术报告:https://arxiv.org/abs/2312.11805(2023-12-19)[✓ 全文亲核] [32] Gemini 2.5 技术报告:https://arxiv.org/abs/2507.06261(2025-07)[✓] [33] Claude Opus 4.5 System Card:https://www-cdn.anthropic.com/bf10f64990cfda0ba858290be7b8cc6317685f47.pdf(2025-11)[✓ 全文亲核] [34] Anthropic《Eval awareness in Claude Opus 4.6’s BrowseComp performance》:https://www.anthropic.com/engineering/eval-awareness-browsecomp(2026-03-06)[✓ 全文亲核] [35] Claude 3 Model Card:https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf(2024-03-04)[✓ 全文检索:无 contamination 章节] [36] Llama 2 论文:https://arxiv.org/abs/2307.09288(2023-07-18)[✓] [37] Llama 3 Herd:https://arxiv.org/abs/2407.21783(2024-07-23)[✓] [38] ConTAM:https://arxiv.org/abs/2411.03923(2024-11)[✓] [39] Sainz et al.《NLP Evaluation in Trouble》:https://aclanthology.org/2023.findings-emnlp.722/(EMNLP 2023 Findings)[✓] [40] 《LLM Benchmark Datasets Should Be Contamination-Resistant》:https://arxiv.org/abs/2605.19999(2026)[✓] [41] Oren et al. 2023:https://arxiv.org/abs/2310.17623(ICLR 2024)[✓] [42] ConStat:https://proceedings.neurips.cc/paper_files/paper/2024/file/a7f89793b9e6f8c6568dbbb6ff727b9b-Paper-Conference.pdf(NeurIPS 2024)[✓] [43] 《Does Data Contamination Detection Work (Well) for LLMs?》(ACL 2025 Findings)[✓ 摘要级] [44] 《Are LLM Benchmarks Already Contaminated? A Systematic Review》(ACL 2026 GEM)[✓ 摘要级] [45] 《Test of Time》:https://arxiv.org/abs/2509.00072(2025-09)[✓ 摘要级] [46] GSM1k:https://arxiv.org/abs/2405.00332(2024-05-01,NeurIPS 2024 D&B)[✓ 全文亲核] [47] TS-Guessing《Investigating Data Contamination》:https://arxiv.org/abs/2311.09783(NAACL 2024)[✓ 全文亲核] [48] Gupta et al.《Changing Answer Order Can Decrease MMLU Accuracy》:https://arxiv.org/abs/2406.19470(ICLR 2025)[✓ 全文亲核] [49] MMLU-CF:https://arxiv.org/abs/2412.15194(ACL 2025)[✓] [50] 《None of the Others》:https://arxiv.org/abs/2502.12896(2025-02)[✓ 摘要级] [51] GSM-Symbolic:https://arxiv.org/abs/2410.05229(2024-10-07,ICLR 2025)[✓] [52] CUTOFF? Beyond?:https://iclr.cc(ICLR 2024)[✓ 摘要级] [53] Leech et al.《Questionable practices in machine learning》:https://arxiv.org/abs/2407.12220(2024-07-17)[✓] [54] Hardt《Test set reuse》:https://mlbenchmarks.org/pdf/05-test-set-reuse.pdf](https://mlbenchmarks.org/pdf/05-test-set-reuse.pdf) [✓ 全文亲核] [55] Dwork et al. 2015《The reusable holdout》:https://pubmed.ncbi.nlm.nih.gov/26250683/(Science)[✓ 摘要] [56] Recht et al. 2019《Do ImageNet Classifiers Generalize to ImageNet?》:https://proceedings.mlr.press/v97/recht19a.html(ICML 2019)[✓ 全文亲核] [57] Schaeffer《Pretraining on the Test Set Is All You Need》:https://arxiv.org/abs/2309.08632(2023-09)[✓] [58] 《Rethinking Benchmark and Contamination》:https://arxiv.org/abs/2311.04850(2023-11-08)[✓ 全文亲核] [59] EvoEval:https://arxiv.org/abs/2403.19114(2024-03-28)[✓] [60] LBPP(EMNLP 2024 Findings)[✓ 摘要级] [61] Leak, Cheat, Repeat:https://arxiv.org/abs/2402.03927(EACL 2024)[✓ 摘要级] [62] SWE-Bench+:https://arxiv.org/abs/2410.06992(2024-10-09)[✓] [63] LessLeak-Bench:https://arxiv.org/abs/2502.06215(2025-02-10)[✓] [64] The SWE-Bench Illusion:https://arxiv.org/abs/2506.12286(2025-06)[✓] [65] Time Travel in LLMs:https://arxiv.org/abs/2308.08493(2023-08-16,ICLR 2024 Spotlight)[✓] [66] 《When Benchmarks are Targets》:https://aclanthology.org/2024.acl-long.744/(ACL 2024)[✓ 全文亲核] [67] The Leaderboard Illusion:https://arxiv.org/abs/2504.20879(2025-04-29)[✓ 全文亲核] [68] LMArena 回应:https://arena.ai/blog/our-response(2025-05-09)[✓] [69] Arena Leaderboard Changelog:https://arena.ai/blog/leaderboard-changelog/(2025-07 起)[✓] [70] OpenAI o1 发布博客:https://openai.com/index/learning-to-reason-with-llms/(2024-09-12)[✓] [71] Bubeck et al.《Sparks of AGI》:https://arxiv.org/abs/2303.12712(2023-03-22)[✓] [72] Stanford AI Index 2025:https://hai.stanford.edu/ai-index/2025-ai-index-report(2025-04-07)[✓] [73] DeepSeek-R1 技术报告:https://arxiv.org/abs/2501.12948(2025-01)[✓] [74] Computerworld 转引 DeepSeek AIME:https://www.computerworld.com/article/3808579/(2025-01-23)[✓] [75] VentureBeat DeepSeek MMLU:https://venturebeat.com/(2025-01-20,报道页对本环境 429,浏览器可访问;本条为二手报道)[◐] [76] CNN 百度文心:https://www.cnn.com/2023/10/17/tech/china-baidu-ernie-ai-upgrade-intl-hnk(2023-10-17)[✓] [77] CNBC:https://www.cnbc.com/2024/03/31/in-ai-race-with-us-china-is-behind-on-a-key-weapon-its-own-openai.html(2024-03-31)[✓] [78] Messick 1995:https://files.eric.ed.gov/fulltext/ED380496.pdf(American Psychologist 50(9):741-749)[✓] [79] 《Standards for Educational and Psychological Testing》2014(AERA/APA/NCME)[◐ 关键页亲核] [80] ETS Upholding Integrity:https://www.ets.org/news/stories/upholding-integrity-ets-unwavering-comittment-test-security.html](https://www.ets.org/news/stories/upholding-integrity-ets-unwavering-comittment-test-security.html) [✓] [81] ETS GRE Test Security:https://www.ets.org/gre/score-users/about/test-security.html](https://www.ets.org/gre/score-users/about/test-security.html) [✓] [82] ETS Why and How ETS Questions Test Scores:https://www.ets.org/pdfs/about/why-how-questions.pdf](https://www.ets.org/pdfs/about/why-how-questions.pdf) [✓] [83] Popham 2001《Teaching to the Test?》:https://www.ascd.org/el/articles/teaching-to-the-test(Educational Leadership 58(6))[✓] [84] Briggs 2009:https://files.eric.ed.gov/fulltext/ED505529.pdf(NACAC)[✓ 全文亲核] [85] 《Measurement to Meaning》:https://arxiv.org/abs/2505.10573(2025-05)[✓] [86] 《Measuring what Matters》:https://proceedings.neurips.cc/paper_files/paper/2025/file/1967e0fc3aa6cbbace562f5cb8e3954e-Paper-Datasets_and_Benchmarks_Track.pdf(NeurIPS 2025 D&B)[✓] [87] 《Evaluating General-Purpose AI with Psychometrics》:https://arxiv.org/abs/2310.16379(2023-10-25;CACM 2026-04)[✓] [88] Epoch AI《A Rosetta Stone for AI benchmarks》:https://epoch.ai/publications/a-rosetta-stone-for-ai-benchmarks(2025-10)[✓] [89] Epoch AI《Are AI benchmarks doomed?》:https://epochai.substack.com/p/are-ai-benchmarks-doomed(2026-05-01)[✓] [90] Epoch AI 评测方法论页:https://epoch.ai/benchmarks/about](https://epoch.ai/benchmarks/about) [✓] [91] Bender et al. 2021《On the Dangers of Stochastic Parrots》:https://doi.org/10.1145/3442188.3445922(FAccT 2021)[✓] [92] Raji et al. 2021《AI and the Everything in the Whole Wide World Benchmark》:https://arxiv.org/abs/2111.15366(NeurIPS 2021 D&B)[✓] [93] Carlini et al. 2022《Quantifying Memorization》:https://arxiv.org/abs/2202.07646(ICLR 2023)[✓] [94] Morris et al. 2025《How much do language models memorize?》:https://arxiv.org/abs/2505.24832(2025-05-30)[✓ 摘要级] [95] Jacovi et al.《Stop Uploading Test Data in Plain Text》:https://arxiv.org/abs/2305.10160(2023-05)[✓ 摘要级] [96] 404 Media《OpenAI Training Bot Crawls ‘World’s Lamest Content Farm’》:https://www.404media.co/(2024-04-12)[✓ 用于背景,与 benchmark 污染无关]
核验统计:96 条编号来源。主笔亲核全文/关键段:约 25 条(含 2026 年三份官方文件全文、Opus 4.5 系统卡、GPT-4/Gemini 1.0/Llama 2-3/GSM1k/Li/Gupta/MMLU-Pro/Leaderboard Illusion/WIMBD/Recht/Hardt/Math-Verify/OLL 退役/ARC-AGI-2 公告/TS-Guessing/Rethinking)。摘要级或转引:约 15 条,均已标注。记忆纠错登记:9 条(10.1)。证据截止:2026-08-01。