跳至正文

被全行业当客观数用的 10.0 分

目录

机制裁决第 121 篇 · 对称双向第 116 篇 · 元尺(section E) · 全库第 179 篇 本篇审的是一把被全世界当客观数用的尺子:Common Vulnerability Scoring System(CVSS)。先读三句红线:

  1. 本篇不裁决任何公司、研究者或安全团队的实务对错,不给「该修哪个漏洞」的具体建议,也不对任何单个漏洞的真实风险做判断——只审「CVSS 分数作为一把尺子,承重承不承得住它被赋予的用途」。
  2. 「CVSS 分数反映严重性」与「CVSS 分数等于风险」是两件不同的事:前者是官方口径,后者是把严重性当风险、把技术属性当真实利用概率的升格。本篇对称审计两个方向。
  3. 全篇承重句均给出可点击来源;英文逐字引用一律标注「一手逐字」(全文/页面取回)或「摘要逐字」(仅摘要取回),未取回原文的明确降档。本篇的独家分母账(KEV 全量 1,656 条 × NVD 分数)全部由本篇自算,算法与方法写在第六节,可复现。

零、一句话裁决

CVSS 是一把真实运转、被全世界采用的严重性尺子——真的不是「一个分数能告诉你该先修哪个漏洞」和「9.8 比 9.6 更危险」这两层被声称的定论。 分数被算出来是真实的,分数被当成风险是升格。本篇要拆的,是「严重性」到「风险」之间被跳掉的那几步,以及每个数字后面被省略的分母。

一、本篇测什么:三跳升格与五件可判定的事

本篇审的不是「CVSS 好不好用」,而是「『CVSS 分数』这个数从被算出来到被用于决策,中间经历了哪些跳」。拆成五件可判定的事:

  1. 尺子真不真:CVSS 是什么、谁在维护、版本怎么演进、分数怎么被算出来(守真锚——FIRST 官方规范、版本史、NVD 采用)——这是守真锚。
  2. 名号真不真:「严重性(severity)」与「风险(risk)」是不是同一个词(官方定义、FAQ 逐字、威胁与环境的责任划分)——这是名号账。
  3. 刻度真不真:同一漏洞不同分析师打出的分数是否一致、不同数据库之间是否一致、分数的精度是否有意义(评分者一致性研究、NVD 内部不一致、2026 年 NVD 官方承认算错)——这是刻度账。
  4. 预测真不真:分数高是否意味着更可能被真实利用(CISA KEV 全量与 NVD 分数的交叉、EPSS 对照、被利用但低分的反例)——这是预测账。
  5. 消费真不真:分数在 PCI DSS、联邦政策、保险与补丁优先级里被怎样使用(PCI DSS 6.3.1 条文、BOD 22-01 与 BOD 26-04、CISA 官方 FAQ)——这是消费账。

结构胎记:名号跳 × 刻度跳 × 存在跳,附一条反向红跳。

  • 名号跳:把「严重性(severity)——对漏洞固有属性的技术描述」读成「风险(risk)——需要考虑真实利用概率与业务影响的东西」。官方自己说 Base 分数不是风险,但整个行业的默认用法是拿它当风险排序。
  • 刻度跳:把「分析师主观选择的离散指标(AV/AC/PR/UI 等,每项只有几个档位)加权后的产物」读成「0.0 到 10.0 的连续精确刻度,9.8 与 9.6 的差异值得排序」。分数从离散输入产生,却在输出端被当成精确连续量。
  • 存在跳:把「被打了分的漏洞」读成「所有漏洞/风险高的漏洞」——大量漏洞根本没有分数(NVD 2026 年 4 月起不再对所有 CVE 打分),被打了分的集合本身有偏。
  • 反向红跳:尺子有噪声≠尺子无用(见第八节反虚无账)。

共用动作:把一句关于「这个漏洞的技术属性」的话,读成一句关于「我的环境里这个漏洞有多危险」的话。

与库内相关篇的分界:

  • 与计量学篇(2026-08-02,section E)分界写死:那篇是「尺子被做对」的正面范本(SI 定义常数、GUM 不确定度、KCDB 比对);本篇是「尺子被造出来但被当客观数用」的造尺人问题。两篇互为镜像:计量学问「尺子做对需要什么」,本篇问「尺子做错时错在哪一步」。
  • 与 AI 评测榜篇(2026-08-01)分界:那篇审「榜分=能力」的升格链;本篇审「严重性分数=风险」的升格链——共享同一母题:被读成判断的那个数,本来是关于被测物固有属性的一句话
  • 与过度推销检测器篇(2026-07-20,section E 收官)分界:那篇是方法学封顶;本篇是它的一把新尺子的专题审计——「成熟度承重光谱」直接套用(CVSS 处在「框架内硬」到「半硬机制」之间,取决于用途)。
  • 与碳核算篇(2026-08-02)分界:那篇的尺子是「判据在场而口径缺席」;本篇的尺子是「刻度在场而预测缺席」——两种缺席模式不同。

编号落位:机制裁决第 121 篇·对称双向第 116 篇·全库第 179 篇(2026-08-03 全库扫描确认未被占用)。去重实测:CVSS漏洞评分EPSSseverity scoreCWE 全库零命中CVE- 6 处全是加密货币篇的具体漏洞个案(非评分对象);KEV/CNA 命中全是英文误匹配——处女地

二、守真锚:CVSS 是真实运转的规范,版本史可逐条核数

先清点尺子本身。本库原则:审「声称」必须回到声称的原件

2.1 官方定义:severity 是官方口径的第一关键词

CVSS v4.0 规范文档(Document Version 1.2)开头:

“The Common Vulnerability Scoring System (CVSS) is an open framework for communicating the characteristics and severity of software vulnerabilities.” [一手逐字]

CVSS v4.0 规范文档(全文取回,117KB HTML)

规范同时写明四个度量组与分数的性质:

“CVSS consists of four metric groups: Base, Threat, Environmental, and Supplemental. The Base group represents the intrinsic qualities of a vulnerability that are constant over time and across user environments… Base metric values are combined with default values that assume the highest severity for Threat and Environmental metrics to produce a score ranging from 0 to 10.” [一手逐字]

关键句:「默认值假设威胁与环境指标取最高严重性」——Base 分数不是「这个漏洞现在有多危险」,而是「假设最坏情况下这个漏洞有多严重」。这一定义在整个评估流程里贯穿:

“The Base Score reflects the severity of a vulnerability according to its intrinsic characteristics which are constant over time and assumes the reasonable worst-case impact across different deployed environments.” [一手逐字]

以及打分者是谁:

“Generally, the Base metrics are specified by vulnerability bulletin analysts, product vendors, or application vendors because they typically possess the most accurate information about the characteristics of a vulnerability. The Threat and Environmental metrics are specified by consumer organizations…” [一手逐字]

责任划分是官方的:Base(严重性)由供给方(厂商/公告分析员)打;Threat(威胁,随时间变化)与 Environmental(环境)由消费方(用户组织)打。NVD 只提供 Base(CVSS-B):

“Assessment providers such as product maintainers and other public/private entities such as the National Vulnerability Database (NVD) typically provide only the Base Scores enumerated as CVSS-B.” [一手逐字]

2.2 版本史:2005→2023,五版演进可逐条核数

FIRST 2023-11-01 v4.0 官方新闻稿(页面取回)逐字给出完整版本史:

“Prior to 2005, custom, incompatible rating systems were used to define severity… CVSS version 1 was released in February 2005, developed then by a small group of pioneers with the aim of industry-wide adoption, with FIRST appointed that April to drive the future development… Over a dozen FIRST members of the CVSS Special Interest Group (SIG) collaborated extensively through 2006 and 2007 to revise and improve CVSS version 1 by testing and re-testing hundreds of real-world vulnerabilities releasing version 2 in June 2007. A third version developed the tool further in 2015 introducing the concept of ‘Scope’… Finally, version 3.1 was released in June 2019 which clarified and improved upon version 3.0 without introducing new metrics or values.” [一手逐字]

版本时间线(官方新闻稿逐字):

版本 时间 官方描述
v1 2005-02 「一小群先驱者开发,目标是全行业采用」
v2 2007-06 SIG 十余名成员用数百个真实漏洞反复测试后修订
v3.0 2015 引入 Scope 概念(跨组件影响)
v3.1 2019-06 澄清改进,无新指标
v4.0 2023-11-01 新命名法(CVSS-B/BT/BE/BTE)、补充指标、OT/ICS/IoT 安全指标

CVSS v2 历史文档(2007 年页面)补充:SIG 从 2005 年 4 月开始 v2 工作,「scored all vulnerabilities in the NIST database to understand and compare against v1 to ensure the fidelity of the scoring increased」[一手逐字]。

2.3 NVD 采用:从 2007 年至今的官方数据库

NVD(National Vulnerability Database) 是美国的官方漏洞数据库,自 v2.0 起采用 CVSS(2007-06-20 NVD 2.0 支持 CVSS v2)。截至本篇取证日(2026-08-03),NVD API 返回全库 CVE 总数 372,520 个[一手,API 直查]。

NVD 官方分档表(CVSS 页面)[一手逐字]:

  • v2.0:Low 0.0-3.9 / Medium 4.0-6.9 / High 7.0-10.0(无 Critical 档)
  • v3.x:Low 0.1-3.9 / Medium 4.0-6.9 / High 7.0-8.9 / Critical 9.0-10.0 / None 0.0
  • v4.0:与 v3.x 同档位(None 0.0 / Low 0.1-3.9 / Medium 4.0-6.9 / High 7.0-8.9 / Critical 9.0-10.0)

注意这个版本会计差异:同一个漏洞在 v2 时代最高档是「High 7.0-10.0」,在 v3 以后才拆出「Critical 9.0-10.0」。分数跨版本不可直接比较——这一点在第五节会再次出现。

三、机制账:分数是怎么被算出来的

3.1 打分流程:离散指标 → 公式 → 连续分数

CVSS v3.1 的 Base Score 由八个指标决定(v3.1 规范):Attack Vector(AV: N/A/L/P)、Attack Complexity(AC: L/H)、Privileges Required(PR: N/L/H)、User Interaction(UI: N/R)、Scope(S: U/C)、Confidentiality/Integrity/Availability(C/I/A: H/L/N)。每项只有两到四个离散档位,分析师从中选择,然后套用固定公式算出 0.0-10.0 的连续分数。

v4.0 进一步把 AV 拆出 Attack Requirements(AT)、把 Scope 换成 Vulnerable/Subsequent System 两套影响指标、UI 拆成 Passive/Active——指标更多、公式更复杂,但「分析师主观选择离散档位」的本质不变

3.2 关键事实:分数来自专家排序,不是来自数据校准

v4.0 FAQ 逐字承认分数的产生方式:

“The method for determining the numeric score is new in CVSS v4.0. The number results from expert ranking done by the CVSS SIG members. This is unique in relation to the algebraic formula of weighting metrics in CVSS v3.0 and v3.1.” [一手逐字]

CVSS v4.0 FAQ(全文取回)

「专家排序」——即 v4.0 的分数不是从公式算出来的,而是 SIG 专家对向量字符串两两比较后排序得出的。EPSS 团队(同为 FIRST 旗下)在 EPSS FAQ 中对此有更尖锐的描述:

“Scores are derived from a collection of experts comparing sets of vector strings and identifying the more-badder one without any empirical validation against any real world” [一手逐字]

以及:

“CVSS Base is an ordinal ranking produced by expert judgment with no empirical calibration against outcomes.” [一手逐字]

这是本库审所有尺子时的核心问题:一把尺子的刻度必须有外部校准(与它声称要预测的东西比较过)才有预测意义。CVSS 的刻度来自专家内部一致性(排序一致性),从未与「真实利用」这个它被用于预测的目标做过校准——直到 EPSS 出现。

3.3 谁在打分:CNA 网络

CVE 编号由 CVE.org(MITRE 运营)管理,打分由 CNA(CVE Numbering Authority,编号机构)完成——厂商对自己产品的漏洞打分,NVD 对这些记录进行 enrich。2026-04-15 起 NVD 改变运营模式(见第五节),不再对所有 CVE 打分。

四、名号账:severity 与 risk 之间的那道官方划线

4.1 官方自己说:Base 分数不是风险

v4.0 FAQ「How does CVSS provide input to patch priority?」逐字:

“CVSS-B Base scores are not risk, and should not be used alone for patch prioritization. However, the technical severity assessments that make up CVSS provide inputs to frameworks, such as BOD 26-04, to help guide patch prioritization as part of enterprise risk assessment.” [一手逐字]

CVSS v4.0 FAQ

同一 FAQ 还写明:

“While the CVSS numeric score is a useful shorthand for vulnerability severity, the score itself does not describe the important context that can be conveyed as part of the entire vector string.” [一手逐字]

以及消费者责任:

“It is the responsibility of the consumer to apply Threat and Environmental data into the assessment of the vulnerabilities (preferably using automation) to reduce the scores of the vulnerabilities that are not as important as others.” [一手逐字]

官方划线的完整逻辑链:Base 分数 = 固有严重性(假设最坏情况、假设威胁全被利用)→ 官方要求消费者自己加 Threat/Environmental 修正 → 官方要求不能只用 Base 分数做补丁优先级 → 官方提供 EPSS(预测利用概率)作为补充。整条链上没有一个环节支持「9.8 分 = 该先修」——那是 Base 分数被单独拿出来用时的默认行为。

4.2 名号跳发生在哪一步

官方规范(v2 时代,v2 指南)的原话是:

“Currently, IT management must identify and assess vulnerabilities across many disparate hardware and software platforms. They need to prioritize these vulnerabilities and remediate those that pose the greatest risk.” [一手逐字]

v2 时代造尺人自己就把「补风险最大的」作为目标写进了引言——但 v4.0 时代官方 FAQ 又明说 Base 分数不是风险。二十年间,官方口径从「帮你补风险最大的」收缩到「请你自己加威胁和环境修正」——名号跳的根源在官方文档自己的措辞演进里:最初承诺的是风险排序,后来交付的是严重性描述,中间没有一版文档明确回收过这个承诺。

4.3 名号跳的机制解剖

把「严重性」读成「风险」需要三步,每步都丢失信息:

  1. 丢弃概率:严重性描述的是「如果被利用,多严重」;风险需要「被利用的概率」。Base 分数里没有概率项(v4.0 的 Threat/Environmental 才有,且官方说那是消费者的责任)。
  2. 丢弃环境:同一个漏洞在暴露的公网服务器和隔离的内网机器上风险不同;Base 分数对所有环境相同(官方明说「constant over time and across user environments」)。
  3. 丢弃业务:一个非关键系统的 9.8 分与核心系统的 7.0 分,风险排序未必是前者优先;分数里没有业务价值项(官方明说「regulatory requirements, number of customers impacted… These factors are outside the scope of CVSS」[一手逐字])。

名号跳的本质:把「这一对(漏洞,环境)的风险」读成「这个漏洞的严重性」——把关系的属性读成实体的属性。这与本库数字孪生篇 FDA 指南的措辞(credibility 是「模型+这一个问题」这一对的属性,不是模型的属性)是同一个结构。

4.4 造尺人自己反对合并

v4.0 FAQ 在回答 AIVSS(AI 漏洞评分)问题时逐字写道:

“AIVSS merges software quality, ethics, privacy, and cybersecurity issues into one-size-fits-all risk measurement. Averaging different dimensions creates dangerously misleading perceptions.” [一手逐字]

造尺人反对把不同维度平均成一个数——但 CVSS 自己就是把八到十七个指标压成一个 0-10 的分数。 这句话用在 CVSS 自己身上同样成立,官方没有对这一点做过澄清。

五、刻度账:同一漏洞,不同的人打出不同的分

5.1 Wunder et al. 2024(IEEE S&P):68% 的评估者对自己重复打分都不一致

Wunder, Kurtz, Eichenmüller, Gassmann, “Shedding Light on CVSS Scoring Inconsistencies: A User-Centric Study on Evaluating Widespread Security Vulnerabilities”, IEEE S&P 2024(arXiv:2308.15259,全文取回)是本篇刻度账的核心实证。

摘要逐字:

“We systematically investigate these questions in an online survey with 196 CVSS users. We show that specific CVSS metrics are inconsistently evaluated for widespread vulnerability types, including Top 3 vulnerabilities from the ‘2022 CWE Top 25 Most Dangerous Software Weaknesses’ list. In a follow-up survey with 59 participants, we found that for the same vulnerabilities from the main study, 68% of these users gave different severity ratings. Our study reveals that most evaluators are aware of the problematic aspects of CVSS, but they still see CVSS as a useful tool for vulnerability assessment.” [一手逐字]

关键发现(全文逐字):

  • 跨评估者不一致:对 Stored XSS 的 User Interaction 指标,Finn’s 系数仅 0.0206(几乎无一致性,p=0.462);Reflected XSS 的 UI 指标 76% 选 UI:R 但一致性仍判为 poor(Finn 0.0206-0.0335 区间);Scope 指标对所有漏洞类型 Finn 系数都远低于 0.2(very poor agreement)。
  • 同一人跨时间不一致:9 个月后的 follow-up,68% 的参与者对同一批漏洞给出了不同的严重性评级——同一评估者自己都不稳定。
  • 态度数据:85% 的评估者认为 CVSS 不一致,但 80% 仍认为它是有用的工具;30% 的参与者从未读过 CVSS 文档。
  • P124 引文(全文逐字):

“CVSS is like democracy: the worst system available, except for all the other systems ever tried.” [一手逐字]

5.2 历史一致性研究:Holm 2015、Allodi 2017

Holm, Afridi 2015, “An expert-based investigation of the Common Vulnerability Scoring System”, Computers & Security 53:18-30(Crossref 元数据 + Wunder 论文转述)——专家评估 CVSS v2 分数,38% 的评估与 NVD 原始分数在严重性档位上不同

Wunder 论文对相关工作的梳理逐字:

“Holm and Afridi investigated the accuracy of CVSSv2 scores. The participants of their survey assessed 3 fixed and 7 randomly selected vulnerabilities from NVD. As a result, 38% of evaluations differed in severity from NVD scores.” [一手逐字,转述]

“Allodi et al. measured the accuracy of CVSSv3.0 assessments by evaluators with different IT security knowledge. Security experts, information security students and computer science students evaluated 30 vulnerabilities… The authors concluded that consistency… is low.” [一手逐字,转述]

5.3 NVD 内部的分数不一致:12,866 条「同描述不同分」

Zhang, Cai, Zhang, Zhao, de Carné de Carnavalet, “The Flaw Within: Identifying CVSS Score Discrepancies in the NVD”, IEEE CloudCom 2023(全文 PDF 取回)——发现 NVD 内部存在描述相同或语义相似、分数却相差悬殊的记录:

“Our analysis identified 12,866 entries suffering from such inconsistencies, highlighting the most error-prone Common Vulnerability Scoring System (CVSS) metrics and vulnerability types, as well as the observed score deviation.” [一手逐字]

12,866 条——这是单一数据库内部的不一致,还没算不同数据库之间的。

5.4 2026 年 NVD 官方公告:约 4,500 条 v4.0 分数算错了

NVD 官方新闻 2026-04-28(页面取回)逐字:

“On April 15th, we were alerted to the presence of inaccurate numerical CVSS v4.0 scores for certain CVE records. After analyzing this issue, we determined that, due to an error in how the numerical portion of the CVSS v4.0 scores are calculated and stored, approximately 4,500 CVE records currently have an incorrect numerical score… Only the calculated numerical CVSS v4.0 scores were affected. Approximately 4,500 CVE records (19% of CVE records that have a v4.0 score) were assigned a CVSS numerical score higher than the correctly calculated value. Fewer than 30 CVE records were assigned a score lower than the correctly calculated value.” [一手逐字]

这是本篇刻度账最硬的一格:被全行业当客观数用的分数,其官方数据库在 2026 年 4 月自查出约 4,500 条(占所有 v4.0 评分记录的 19%)数值算错,且错误方向是系统性偏高(只有不到 30 条偏低)。NVD 承诺 2026-04-28 修复并引入自动化验证。

同时NVD 新闻 2026-04-15:NIST 改变运营模式,不再分析所有 CVE,只对符合条件的 CVE 做 enrich(打分数),其余标为 lowest priority 不立即打分。

“In the past, NIST’s NVD program aimed to analyze all CVEs to add details — such as severity scores and product lists… Going forward, NIST will add details, or ‘enrich,’ those CVEs that meet certain criteria… CVEs that do not meet those criteria will still be listed in the NVD but deemed as ‘lowest priority’ and will not be immediately enriched by NIST.” [一手逐字]

存在跳的官方源头:从 2026 年起,「有 CVSS 分数的漏洞」与「全部漏洞」的差距会系统性扩大——NVD 不再给所有 CVE 打分。

5.5 假精度:离散输入如何变成连续输出

CVSS 的输入是离散档位(AV 4 档、AC 2 档、PR 3 档、UI 2 档、S 2 档、C/I/A 各 3 档),但输出是 0.0-10.0 的连续分数。v3.1 的公式会产出 9.8、9.6、9.4 这样的分数。问题在于:输入的任何一档都是分析师判断,而分析师判断的一致性低到 5.1 节所示水平——那么 9.8 与 9.6 的差异,反映的是真实严重性差异,还是不同分析师的判断差异?当评估者自身重评都有 68% 的档位变化时,分数小数点后一位的精度是伪精度。官方 FAQ 承认:

“Why do some unique CVSS v4.0 metrics result in the same numeric score?… Why do some metrics, such as User Interaction or Privileges Required, have less than expected impact on the resulting score?” [一手逐字,问题原文]

六、预测账:分数高,真的更可能被利用吗——本篇自算分母账

这是本篇的独家核心。方法:拿 CISA KEV(已知被利用漏洞目录)全量与 NVD 的 CVSS 分数做交叉,自算「被真实利用的漏洞」与「全部漏洞」的分数分布对比

6.1 方法与数据(可复现)

  1. KEV 全量CISA KEV JSON feed(2026-07-29 快照)——1,656 个 CVE,全部为 CISA 判定有可靠证据在野利用的漏洞。
  2. NVD 分数:NVD API(https://services.nvd.nist.gov/rest/json/cves/2.0)逐条拉取 1,653/1,656 个 KEV CVE 的 CVSS 分数(3 个 2002-2006 年老 CVE 无 API 返回);全库参照样本 20,000 个有分记录 + NVD API 返回全库总数 372,520。
  3. 统计:按 v3.x 档位(与 v2 不可直接比,KEV 中全部有 v3.x 分)分档计数。

6.2 被利用的漏洞偏高分——真

档位 KEV(被利用,n=1,653) 全库(NVD 样本,n=20,000)
Critical 9.0-10 578(35.0%) 1,637(8.2%)
High 7.0-8.9 869(52.6%) 6,948(34.7%)
Medium 4.0-6.9 198(12.0%) 9,668(48.3%)
Low 0.1-3.9 8(0.5%) 1,729(8.6%)

被利用的漏洞确实偏向高分:KEV 中 87.6% 在 7.0 以上,而全库只有 42.9%。「CVSS 高分与真实利用正相关」是真命题——这就是支持侧的证据,本篇不抹掉它。

6.3 但反向账更硬:约 98% 的 Critical 漏洞从未被列入已知利用

全库 Critical(9.0+)占比 8.2% × 全库 372,520 ≈ 30,491 个 Critical 漏洞。KEV 中 Critical 仅 578 个。

这意味着:约 98.1% 的 Critical 漏洞不在 KEV 里——从未被 CISA 判定有在野利用证据。

「Critical」这个标签给出的信号(这个漏洞很危险)与真实利用之间,隔着约 50 倍的漏报。用 Critical 排序修补清单,意味着把大量从未被利用的 9.x 分漏洞排在真正被利用的漏洞前面——这正是 BOD 22-01 FAQ 官方逐字说的:

“Also, many vulnerabilities classified as ‘critical’ are highly complex and have never been seen exploited in the wild – in fact, less than 4% of the total number of CVEs have been publicly exploited.” [一手逐字]

“What is more important to remediate first – critical and high or known exploited vulnerabilities? Known exploited vulnerabilities should be the top priority for remediation.” [一手逐字]

6.4 反例:被利用但低分

KEV 中被利用但 CVSS < 7.0 的有 206 个(12.5%),其中 < 4.0(Low)的 8 个:

CVE 分数 备注
CVE-2025-47729 1.9 已被利用(KEV 收录)
CVE-2024-55550 2.7 已被利用
CVE-2021-25489 3.3 已被利用
CVE-2023-26083 3.3 已被利用
CVE-2021-44168 3.3 已被利用
CVE-2013-2423 3.7 已被利用
CVE-2022-23134 3.7 已被利用
CVE-2023-20867 3.9 已被利用

一个 CVSS 1.9 分的漏洞被真实利用——Low 档,按任何标准排在修补清单末尾,却被攻击者用上了。这与「分数低的漏洞不危险」的默认假设直接冲突。

6.5 EPSS 对照:同团队的替代尺子也承认 CVSS 分数预测力有限

EPSS(Exploit Prediction Scoring System) 是 FIRST 旗下的另一个项目,用机器学习预测「CVE 在未来 30 天内被在野利用的概率」,每天更新全部 CVE 的 0-1 概率与百分位。注意:EPSS 与 CVSS 由同一个组织(FIRST)维护——这是造尺人自己造了第二把尺子来补第一把尺子的洞。

EPSS 论文(arXiv:2302.14172)摘要逐字:

“Unfortunately, existing vulnerability scoring systems are either vendor-specific, proprietary, or are only commercially available. Moreover, these and other prioritization strategies based on vulnerability severity are poor predictors of actual vulnerability exploitation because they do not incorporate new information that might impact the likelihood of exploitation.” [摘要逐字]

「基于严重性的优先化策略是真实利用的糟糕预测器」——这是 FIRST 自家论文对 CVSS 的定位。

EPSS FAQ 更直接:

“CVSS and EPSS measure different things and are empirically uncorrelated. High severity scores are slightly better than random at predicting exploitation activity.” [一手逐字]

“Multiplying EPSS by a CVSS score does not compute probability × severity and is never a good idea. EPSS is a calibrated probability and CVSS Base is an ordinal ranking produced by expert judgment with no empirical calibration against outcomes.” [一手逐字]

「高分略好于随机」——这就是官方团队对「CVSS 分数预测真实利用」能力的全部背书。

本篇用 EPSS 数据做同样交叉(可复现):KEV 全量 1,656 个的 EPSS 分布——23.9% 的 KEV 漏洞 EPSS < 0.1(模型认为未来 30 天被利用概率低于 10%),16.4% < 0.05,2.6% < 0.01。EPSS 预测力确实比 CVSS 强(EPSS and CISA KEV 官方分析:KEV 中约一半 EPSS 接近 0;ScoobyLabs 2026-04-09 独立分析:KEV 漏洞平均 EPSS 0.2253 vs 基线 0.0343,6.57 倍提升),但绝对精度仍低(ScoobyLabs:任何阈值下精度 <1%)——「EPSS works as a filter, not a verdict」。没有一把尺子是判决。

七、消费账:分数在制度里怎么被用

7.1 PCI DSS:从点名到不点名

PCI DSS v3.2.1 Requirement 6.2 逐字点名 CVSS:

“Ensure that all system components and software are protected from known vulnerabilities by installing applicable security patches… using a ranking system, such as CVSS (Common Vulnerability Scoring System)…” [一手逐字,Wayback 存档取回]

PCI DSS v4.0.1 Requirement 6.3.1 不再点名 CVSS官方 PDF,Cloudflare 反爬,经 TrustedSec 权威转引核对):

“Security vulnerabilities are identified and managed as follows: New security vulnerabilities are identified using industry-recognized sources… Vulnerabilities are assigned a risk ranking based on industry best practices and consideration of potential impact. Risk rankings identify, at a minimum, all vulnerabilities considered to be a high-risk or critical to the environment.” [一手逐字,经转引核对]

TrustedSec 对 6.3.1 的解析(QSA 视角):6.3.1 要求「风险排名」但不指定具体系统——CVSS 从「被法规点名的尺子」降级为「可选的排名方法之一」。PCI DSS 4.0 还把 Critical/High 漏洞的补丁时限写进 6.3.3(一个月内),把「哪个漏洞算 Critical/High」这个定义权留给了组织自己——而绝大多数组织会直接用 CVSS 档位(见 TrustedSec 解析)。

7.2 CISA 政策:从 BOD 22-01 到 BOD 26-04

BOD 22-01(2021-11 发布,[官方页面)](https://www.cisa.gov/news-events/directives/bod-22-01-reducing-significant-risk-known-exploited-vulnerabilities)**:建立 KEV 目录,要求联邦机构按 KEV 时间线修补——把「已知被利用」设为修补的第一优先级,明说优先于 critical/high**。FAQ 逐字:

“CVEs are currently scored under the CVSS system, which does not take into consideration whether a vulnerability has ever been used to exploit a system in the wild.” [一手逐字]

“Rather than have agencies focus on thousands of vulnerabilities that may never be used in a real-world attack, BOD 22-01 shifts the focus to those vulnerabilities that are active threats. CISA acknowledges CVSS scoring can still be a part of an organization’s vulnerability management program.” [一手逐字]

BOD 26-04([官方页面)](https://www.cisa.gov/news-events/directives/bod-26-04-prioritizing-security-updates-based-risk)**:2026 年取代 BOD 22-01(官方页面标注 Revoked),改用 SSVC 决策树(KEV 状态 + 利用自动化 + 技术影响)。CVSS 在其中的位置被官方逐字写明**:

“Table 1: Remediation Timelines is informed by the SSVC system… Technical impact depicts how much post-exploitation control an adversary gains over the affected asset and is similar to the Common Vulnerability Scoring System (CVSS) base score’s concept of ‘severity.'” [一手逐字]

「similar to…’severity’」——在最新联邦政策里,CVSS 只作为「技术影响」这一维度的近似概念存在,且由 CISA 通过 Vulnrichment 程序自己为每个 CVE 提供该维度。联邦政策的演进方向是:从「用 CVSS 分数排序」到「用 KEV+SSVC 排序,CVSS 只贡献技术影响一维」。

7.3 保险业:问卷为主,CVSS 只是背景

国际信息安全杂志 2026 综述(Paragioudakis et al.)(Crossref 摘要):

“Traditional cyber insurance underwriting still relies heavily on questionnaire-based risk assessments, lacking a truly data-driven approach.” [摘要逐字]

保险业主要靠问卷(含补丁管理问题),CVSS 分数不是定价的直接输入——「保险用 CVSS 定价」的说法在本篇取证范围内未找到一手证据,如实降级为未确认

7.4 消费账小结

三个最重的消费方(PCI DSS、CISA 政策、保险)中,两个已经或正在把 CVSS 从「主尺」降为「输入之一」:PCI DSS 从点名到不点名,CISA 从依赖分数到 KEV/SSVC 优先。被全行业当客观数用的地位,与官方自己的制度演进方向相反。

八、反虚无账:尺子有噪声,不等于尺子无用

对称审计的另一半:「CVSS 全是废物」同样不成立。

  1. 严重性描述本身是真功能:CVSS 规范、矢量字符串(vector string)提供了一种跨厂商、跨平台沟通漏洞技术属性的标准语言。8 个指标的离散选择承载了真实的技术信息(攻击向量、是否需要权限、影响范围)——这些信息本身是对的,问题只在「压成一个数」之后。Wunder 论文 80% 的评估者仍认为 CVSS 有用,P124 的「democracy」类比是真实写照:没有更好的替代品。
  2. 正相关是真:6.2 节已证,KEV 中 87.6% 分数 ≥7.0——高分数与真实利用存在正相关,作为「最低门槛过滤」(先把 9.x 看完)有一定筛选价值。CISA 官方 FAQ 说 CVSS「can still be a part of an organization’s vulnerability management program」[一手逐字]。
  3. 阈值有制度价值:PCI DSS 6.3.3 的「Critical/High 一个月内补」把分数档位变成合规硬约束——无论分数精度如何,档位作为「至少处理这些」的底线清单是有效的(弱筛)。
  4. EPSS 用 CVSS 指标做输入EPSS 论文的特征包含 CVSS 矢量信息——CVSS 的技术描述是 EPSS 预测的特征来源之一。尺子的部分信息被第二把尺子有效利用了

反虚无的结论:CVSS 的价值在「技术严重性的标准沟通语言」和「粗筛门槛」层面是真实的;它在「精确排序修补优先级」层面的升格是虚假的。两件事必须分开。

九、边界账与诚实空位

  1. 版本不可比:v2 无 Critical 档(High 7.0-10.0),v3/v4 拆出 Critical 9.0-10.0。跨版本比较分数(例如用历史 v2 分做研究)需明确版本。
  2. v4.0 采用率:NVD 截至 2026-04-15 才承认 4,500 条 v4.0 分数算错并修复,v4.0 全面采用仍在进行中(TrustedSec 2024 年仍建议用 v3.1);v4.0 分数的分布与 v3.1 不同(FAQ 承认「Scores for the same vulnerability are different between v3.1 and v4.0」)。
  3. 打分者身份偏置:Base 由厂商/公告分析员打分(官方明说),厂商对自己的产品漏洞打分存在利益相关(压低严重性的动机),本篇未找到系统性实证,如实列为待查。
  4. KEV 本身的偏置:KEV 收录依赖「可靠的在野利用证据」,存在系统性漏报(特别是小厂商产品、闭源软件);「不在 KEV」不等于「从未被利用」。本篇所有「未列入 KEV」的说法均指「未被 CISA 判定」,不指「从未被利用」。
  5. 全库样本偏置:本篇全库参照样本(20,000 个有分记录)来自 NVD API 的分页拉取,覆盖 2023-2026 年 lastModified 的 CVE 为主,与全库 372,520 的构成可能有偏差;Critical 占比 8.2% 的估计已用全库总数交叉,但未做完全抽样。
  6. 三个 2002-2006 年 KEV CVE 无 NVD API 返回(CVE-2002-0367、CVE-2004-0210、CVE-2006-1547),未计入分布。
  7. 保险业 CVSS 使用:未找到一手证据,如实留空(见 7.3)。
  8. NVD 2026-04-15 运营变更的影响:2026 年 4 月后新 CVE 的打分覆盖大幅下降,本库统计截止 2026-07-29(KEV 快照),后续覆盖变化未跟踪。

十、造尺人层:FIRST 自己的三句话

本篇的造尺人(FIRST/CVSS-SIG)在自家文档里留下了三句关键自白,放在一起读:

  1. 官方规范:「assumes the reasonable worst-case impact across different deployed environments」——Base 分数假设最坏情况,不是预测。[一手逐字]
  2. 官方 FAQ:「CVSS-B Base scores are not risk, and should not be used alone for patch prioritization」——官方明说 Base 分数不是风险。[一手逐字]
  3. 官方 FAQ:「The number results from expert ranking done by the CVSS SIG members」——v4.0 分数来自专家排序,无外部校准。[一手逐字]

再加同组织 EPSS 团队的第四句:

  1. EPSS FAQ:「High severity scores are slightly better than random at predicting exploitation activity」——同组织对「CVSS 预测真实利用」的背书上限。[一手逐字]

造尺人把每一条免责都写在了文档里。 被省略的不是免责声明,而是「当行业把 9.8 分当作风险排序依据时,这些免责没有被消费」这一事实。

十一、母题收口:被读成判断的那个数

本篇与 AI 评测榜篇共享母题:被读成判断的那个数,本来是关于被测物固有属性的一句话。

  • AI 榜分(MMLU/SWE-bench)测量的是「模型在测试集上的表现」,被读成「模型的能力」。
  • CVSS 分数测量的是「漏洞的技术严重性(假设最坏情况)」,被读成「这个漏洞在我的环境里有多危险、该先修」。

两次跳的结构相同:把测量的读数当成对被测物价值的判断,把「关于它的一句话」当成「关于我该怎么办的一句话」。

灵魂句:CVSS 分数是工程师写给工程师的「这个漏洞有多坏」的便签,被全行业读成了「先修哪个」的判决书——便签是真的,判决书不是;而造尺人把这两件事的分界线,写在了 FAQ 第 1006 行。


来源清单

  1. CVSS v4.0 规范文档(FIRST 官方,Document Version 1.2,全文取回)——「open framework for communicating the characteristics and severity」「reasonable worst-case」「agnostic to the individual」「Base metrics are specified by vulnerability bulletin analysts」等承重句 [一手逐字]
  2. CVSS v4.0 FAQ(FIRST 官方,全文取回)——「CVSS-B Base scores are not risk」「expert ranking」「provider must assume that all vulnerabilities will be exploited and fully weaponized」「It is the responsibility of the consumer」「None of these scoring systems is a replacement」「BOD 26-04」「AIVSS merges… one-size-fits-all risk measurement」「scores clustered toward Critical and High ratings was not a problem」 [一手逐字]
  3. CVSS v4.0 用户指南(FIRST 官方,全文取回)——评分细则 [一手逐字]
  4. CVSS v4.0 实施指南(FIRST 官方,全文取回)——消费者实施建议 [一手逐字]
  5. CVSS v3.1 规范文档(FIRST 官方,全文取回)——八指标定义 [一手逐字]
  6. CVSS v2 完整文档(FIRST 官方,全文取回)——v2 引言「remediate those that pose the greatest risk」 [一手逐字]
  7. CVSS v1 完整文档(FIRST 官方,全文取回)——2005 年原版 [一手逐字]
  8. CVSS v2 History(FIRST 官方 2007 年页面,全文取回)——v2 开发过程 [一手逐字]
  9. FIRST 2023-11-01 v4.0 官方新闻稿(页面取回)——完整版本史逐字 [一手逐字]
  10. FIRST 2023-11-01 新闻稿 PDF(PDF 取回)——同上内容 PDF 版 [一手逐字]
  11. CVSS-SIG 首页(页面取回)——SIG 使命「numerical score reflecting its severity」 [一手逐字]
  12. EPSS 首页(页面取回)——「estimates the probability that a published CVE will be exploited in the wild in the next 30 days」 [一手逐字]
  13. EPSS FAQ(页面取回)——「CVSS measures severity, but severity is never defined」「empirically uncorrelated」「slightly better than random」「never a good idea to multiply」 [一手逐字]
  14. EPSS Why(页面取回)——「A scoring system that is not grounded in outcome data is simply adding structure to an assertion」「EPSS scores are calibrated probabilities」 [一手逐字]
  15. EPSS 论文(arXiv:2302.14172)(摘要取回)——「prioritization strategies based on vulnerability severity are poor predictors of actual vulnerability exploitation」 [摘要逐字]
  16. EPSS 原论文(arXiv:1908.04856,Jacobs et al.)(摘要取回)——EPSS 原始定义 [摘要逐字]
  17. Wunder et al., IEEE S&P 2024(arXiv:2308.15259)(全文 HTML 取回)——196 人调查、68% 同一人重评不同、Finn 系数、85%/80%/30%、P124 引文、Holm 38% 与 Allodi 转述 [一手逐字]
  18. Holm & Afridi 2015, Computers & Security 53:18-30(Crossref 元数据 + Wunder 转述)——38% 评估与 NVD 不同 [多源交叉]
  19. Zhang et al., IEEE CloudCom 2023(全文 PDF 取回)——NVD 内部 12,866 条不一致 [一手逐字]
  20. CISA KEV 目录(官方页面)——KEV 收录标准 [一手逐字]
  21. CISA KEV JSON feed(2026-07-29 快照,1,656 条)——本篇分母账数据源 [一手逐字]
  22. BOD 22-01 官方页面(页面取回)——「less than 4%」「does not take into consideration whether a vulnerability has ever been used」「Known exploited vulnerabilities should be the top priority」「CISA acknowledges CVSS scoring can still be a part」 [一手逐字]
  23. BOD 26-04 官方页面(页面取回)——SSVC 决策树、「Technical impact… similar to the CVSS base score’s concept of ‘severity’」 [一手逐字]
  24. NVD 新闻页(页面取回)——2026-04-28 v4.0 分数错误公告(4,500 条、19%、系统性偏高)、2026-04-15 运营变更公告 [一手逐字]
  25. NVD API(本篇直查 372,520 全库总数 + 1,653 条 KEV 分数)——本篇自算分母账数据源 [一手逐字]
  26. NVD CVSS 分档表(页面取回)——v2/v3/v4 档位表 [一手逐字]
  27. PCI DSS v4.0.1实测 403,Cloudflare 反爬,条文经 TrustedSec 权威转引双通道核对)——6.3.1 条文 [多源交叉]
  28. TrustedSec: PCI DSS 6.3.1 解析(QSA 视角)(页面取回)——6.3.1 条文转引、CVSS 作为可选排名方法 [多源交叉]
  29. CISA Vulnrichment 仓库(API 索引 79,900 条 + 1,313 条 KEV 文件取回)——CISA 自己补的 CVSS v4 评分 [一手逐字]
  30. EPSS and CISA KEV 官方分析(页面取回)——「Approximately half of CISA KEV CVEs have an EPSS score near 0」 [一手逐字]
  31. ScoobyLabs EPSS Accuracy 分析(2026-04-09)(页面取回)——6.57× lift、精度 <1%、「filter, not a verdict」 [一手逐字]
  32. Paragioudakis et al. 2026, Int. J. Inf. Secur.(DOI 10.1007/s10207-026-01275-5)(Crossref 摘要)——保险业问卷为主 [摘要逐字]
  33. CVE.org(官方入口)——CNA 网络背景 [一手逐字]