目录
- 零、一句话裁决
- 一、本篇测什么
- 二、核验标记与层速览
- 三、守真锚(这一侧一条都不动)
- 四、仪器账:这台机器到底测什么(信号跳·物理层)
- 五、提问账:三种问法,三种理论(信号跳·逻辑层)
- 六、效度账:个案调查里它有多好(信号跳与分母跳的接缝)
- 七、分母账:10,000 人里抓 10 个间谍的算术(本篇最重格)
- 八、间谍账:测谎没有抓到的那几个人
- 九、判例账:669 词如何统治美国证据法七十年
- 十、法域账:拼布、例外与国际
- 十一、名号与消费账:谁叫它「测谎仪」
- 十二、反措施账:约半数能破,教这件事的人入狱
- 十三、反向红跳:四句流行的「反转」也逐句裁决
- 十四、对称金句与三向红线
- 附录 A:任务书预设勘误明细(6 处)
- 附录 B:取不回清单(缺席登记,缺席不等于不存在)
- 附录 C:来源清单(全部为本轮实际取回并落盘件;Wayback 链接为 id_ 原件快照)
机制裁决第 151 篇 · 对称双向第 146 篇 · section H · 全库第 209 篇 本篇审的是「测谎仪能读出谎言/测谎可用于安全筛查/测谎结果被法庭采纳」这组话——以及它背后那台记录呼吸、血压与皮肤电的仪器、那套对照问题话术、那张 10,000 人里抓 10 个间谍的自算表、和那个不足两页的 1923 年判决。先读三句红线:
- 本篇不裁决任何人该不该接受测谎,不构成任何法律、就业或安全建议;涉及间谍案、刑事案件处只审公开档案里测谎环节的角色,不重新评价案件本身。
- 「多道仪如实记录了生理信号」与「这台机器能从人体内读出真话」是两件不同的事。本篇的裁决落在两者之间的缝里:仪器是真的(它记录的曲线是真的),个案判别好于随机是真的(官方委员会逐字承认),而「测谎」这个名号是假的(连发明它的三个人都拒绝这个词)——跳变不发生在仪器里,发生在「唤醒」被读成「谎言」的那个动作里。
- 全篇承重句均给出可点击来源;中英文逐字引用一律取自实际取回并落盘的文件(含 Wayback 原件),自算部分写明算式。本篇引用的所有数字(A 值、假阳性数、测谎次数、引用次数)都给出出处与口径。
零、一句话裁决
「测谎仪测的不是谎」是真的——从业者自己都不声称仪器直接测量欺骗:「Practitioners do not claim that the instrument measures deception directly」(NAS 2003 第 1 章[一手逐字]);欺骗没有特异的生理反应也是真的——这句话美国国会技术评估办公室 1983 年就写了(「There is no known physiological response that is unique to deception」,OTA 1983[一手逐字]),国家科学院 2003 年又写了一遍(「There is no unique physiological response that indicates deception」,NAS 2003 第 3 章[一手逐字])。个案刑事侦查中测谎判别显著好于随机也是真的——同一份 NAS 报告的结论逐字是「well above chance, though well below perfection」(执行摘要[一手逐字])。被读成的一句假话,是把「一个记录情绪唤醒的仪器在受控提问下与欺骗的统计关联」读成「一台能从人体内读出真话的机器」的那个动作——而这个动作最贵的一次,是把个案侦查的效度读成国家安全筛查的执照:10,000 名雇员里藏 10 个间谍,要抓住其中 8 个,就得把 1,598 名清白雇员一起标记为「欺骗者」——这是委员会自己算的表(NAS 2003 表 S-1[一手逐字]),而做出这份裁决二十多年后,联邦政府的测谎量不减反增至每年十一万次以上。
本篇的独占格在第七章(分母账:NAS 自算表逐字还原+假阳性指数账+DOE 2006 年规则这个罕见的「报告落地」样本)、第八章(间谍账:Ames 三次测谎、Hanssen 二十五年从未被测、Montes「passed a polygraph」——测谎史上三宗最著名的间谍,仪器在他们身上的真实表现)、第九章(判例账:Frye 1923 全文 669 词、沉睡五十年的引用计量、被 Daubert 取代、Scheffer 8-1——产出美国证据法总判例的那台机器,自己至今进不了大多数法庭)、第十二章(反措施账:约半数受训者能击败测试,而教这件事的人被判了两年——反措施产业与 Operation Lie Busters)。
灵魂句:机器记录的是唤醒,宣判的是人。一百年来被测试的从来不是谎言——是一个制度愿意为多低的基数、错标多少清白者。那台机器给美国法律留下了统治科学证据七十年的判例,而它自己,至今不被允许出庭作证。
一、本篇测什么
1.1 被审对象
被审的是一组互相咬合的说法及其用法:
- 信号句:「说谎会引起特有的生理反应,仪器能把它读出来」(从仪器宣称到大众常识)
- 分母句:「测谎可用于雇员/安全筛查」(从个案效度到国家制度的升格)
- 判例句:「测谎证据被(或不被)法庭采纳」(从一页纸意见到七十年总判据的一生)
- 名号句:「lie detector/测谎仪」(从报界造词到行业招牌)
1.2 结构胎记
信号跳 × 分母跳 × 判例跳 + 反向红跳。
- 信号跳:把「交感神经唤醒指标(呼吸、血压、皮肤电)的变化」读成「欺骗的特异性生理签名」——NAS 2003 与 OTA 1983 相隔二十年的同一句否定;1917 年 12 月 14 日 Shepard 致 Yerkes 的信在该技术诞生之初就写下同款批评(「something doing, but not of any particular kind of activity」[一手逐字]);三种提问技术(RIT/CQT/CIT)各配一套理论,而委员会对研究的总体定性是「atheoretical」[一手逐字]。
- 分母跳:把「个案刑事侦查的定向效度」读成「大规模雇员/安全筛查的可用工具」——NAS 2003 自算表(表 S-1/表 2-1)逐字;假阳性指数与 PPV 账(刑事场景 81% vs 筛查场景 0.4%);「generalization from them to uses for screening is not justified」直接句;Ames/Hanssen/Montes 实证链;EPPA 1988 的豁免结构——私人雇主被禁、政府被豁免于自己签署的法律之外,效度最弱的场景恰是法定保留的场景。
- 判例跳:把「1923 年 12 月 3 日一份不足两页、未引用任何先例的特区上诉意见」读成「此后七十年美国科学证据可采性的总判据」——Frye 全文 669 词逐字;沉睡史量化(判决后十年零引用;前 25 年联邦 8 次+州 5 次;1980 年代每年引用量≈前 50 年总和);Daubert 1993 取代句逐字;Scheffer 1998 最高法院 8-1 维持军事法庭排除;讽刺格=产生该判例的那台装置本身,至今在大多数法庭不可采。
- 反向红跳(第十三章):四句反向审计——「测谎和抛硬币一样」不立(好于随机是委员会逐字结论);「反措施傻瓜可破」部分为真但被夸大(约半数成功率、CIT 效应混杂、「能否骗过训练有素的检查员」官方答 unknown);「美国法律全禁」不立(EPPA 只禁私人雇主;30/17/12 的州格局;新墨西哥常态可采);「脑成像已解决测谎」不立(Semrau 2012 排除+「no region could be used to correctly detect deception across all individuals」[一手逐字])。
1.3 落位:section H 与分界
本篇落 section H(社会/制度),与三篇邻近篇分界写死:
- 与司法鉴定篇(2026-07-31,第 111 篇):那篇审痕迹比对的个体化宣称(指纹/枪弹/咬痕/毛发——对象=物理痕迹的同一认定),本篇审欺骗检测(对象=生理信号到心理状态的推断)。两篇共享 Frye/Daubert/Rule 702 法律框架,但分工明确:那篇的判例主角是 Daubert/PCAST/NAS 2009,本篇的判例主角是 Frye 本身——NAS 2003《The Polygraph and Lie Detection》与 NAS 2009《Strengthening Forensic Science》是两份不同报告,正好构成国家科学院自我否定的姊妹卷:2003 年判测谎筛查不可用,2009 年判多数法庭学科基础不足。
- 与 QKD 篇(2026-08-13,第 150 篇)同族不同物:那篇把「原理安全」读成「工程安全」,本篇把「唤醒信号」读成「欺骗签名」——共用「限定词印在文件里、错位发生在文件外」的形状。
- 与司法鉴定篇、CVSS 篇、教育测量篇同属「分母与尺子」元尺家族:本篇是其中唯一一件「造尺者(委员会)亲手写下分母算术,而制度照用不误」的样本。
1.4 去重实测
python 全库扫描 208 篇 / 7,720,699 字符:测谎 仅 1 处(同行评审篇「它不是测谎仪」比喻义);polygraph/lie detector/Frye/Lykken/Iacono/CQT/GKT/CIT/对照问题/CVSA/No Lie MRI 全库零命中;Daubert 6 处全在司法鉴定篇(本篇只衔接不重复);皮肤电 2 处为过拟合篇机制义——测谎专篇零命中,处女地确认。
1.5 边界
- 不裁决任何在押或已决案件的具体对错;间谍案只审「测谎环节在其中扮演了什么角色」这一层。
- 不提供任何「如何通过/应对测谎」的操作信息;反措施一章只到「实证研究显示什么、官方立场是什么、法律如何处置教反措施的人」为止。
- 中国法域只登记官方规范文本与官方行业报话语,不评价其司法实践。
- 本篇不涉及任何测谎服务的选购建议。
1.6 方法备案
- 取证通路:NAP 逐章阅读页
nap.nationalacademies.org/read/10420/chapter/N(必须--compressed,否则 0 字节);CourtListener v4 检索 API 匿名可用、判决全文走 Wayback CDX+id_原件;Cornell LII 直取;govinfo 法典与听证记录直取;联邦公报 API(curl-g);Europe PMC/PubMed efetch;polygraph.org 官方文档库直取;DCSA 预算件走 comptroller.war.gov;dcsa.mil/oig.justice.gov 旧链 403/404 走 Wayback。 - 承重统计一律 python3 复算;NAS 表格数字按页面内嵌页码锚点定位。
- 任务书预设勘误 6 处(全部有落盘证据,明细见附录 A):「2013 年最高法刑诉法解释排除心理测试」表述证伪(三个同期规范文本全文零命中,实际排除依据是 1999 年最高检批复+证据种类封闭清单);Lee v. Martinez 正确案号为 2004-NMSC-027(任务书预设 051);NAS 自算表用 A=0.90 作乐观上界(池表写「AUC 0.85」,0.85 级数字对应实测中位 0.86);「Marston 是 Frye 案专家证人」不准确(测了、被提名、被拒作证、判决不点名);Frye「一页纸」表述按 669 词/不足两页修正;Iacono & Lykken 1997 精确百分比原文付费墙未破,本篇只承接到「约三分之一/约四分之一」级(双源交叉)。
二、核验标记与层速览
核验标记沿用本库四档:[一手逐字]=原文逐字取自实际取回文件(含 Wayback id_ 原件);[文献较稳]=多源一致或权威评估;[需亲核]=正文未全取回或版本存疑;[多源检索]=方法学阴性检索(含「缺席即证据」登记)。
| 层 | 名称 | 核心问题 | 裁决方向 |
|---|---|---|---|
| 三 | 守真锚 | 仪器/好于随机/CIT/规模 | 一条不动 |
| 四 | 仪器账 | 机器到底测什么 | 唤醒,不是谎言 |
| 五 | 提问账 | 三种问法三种理论 | 话术承重,理论缺席 |
| 六 | 效度账 | 个案里它有多好 | 好于随机、远低于完美、系统性高估 |
| 七 | 分母账 | 筛查算术 | 官方自算表判不可用,制度照用 |
| 八 | 间谍账 | 实证层 | 三名著名间谍全部穿过筛查体系 |
| 九 | 判例账 | Frye 的一生 | 669 词沉睡五十年,统治七十年 |
| 十 | 法域账 | 州拼布与国际 | 拼布;中国可辅不可证;日本只用 CIT |
| 十一 | 名号与消费账 | 谁叫它 lie detector | 报界造词,发明者拒用 |
| 十二 | 反措施账 | 攻与防 | 约半数可破;教人防者入狱 |
| 十三 | 反向红跳 | 四句反向读法 | 三句不立,一句部分为真被夸大 |
三、守真锚(这一侧一条都不动)
在拆任何跳变之前,先把真实的那一侧钉死。以下六条本篇不做任何弱化:
守真锚一:仪器真实,记录真实。 多道生理记录仪如实记录呼吸、相对血压与皮肤电——「a piece of equipment that records physiological phenomena—typically, respiration, heart rate, blood pressure, and electrodermal response」(NAS 2003 第 1 章[一手逐字])。仪器本身不撒谎;争议从不在「曲线是不是那天的曲线」。
守真锚二:个案判别显著好于随机——这是官方委员会的结论,不是行业的广告。 「specific-incident polygraph tests can discriminate lying from truth telling at rates well above chance, though well below perfection」(NAS 2003 执行摘要[一手逐字]);「features of polygraph charts and the judgments made from them are correlated with deception」(第 5 章[一手逐字])。更早的 OTA 1983 结论同向:「the polygraph detects deception at a rate better than chance, but with error rates that could be considered significant」(OTA 1983 第 7 章[一手逐字,扫描 OCR 件])。
守真锚三:CIT(隐蔽信息测试)有真实的科学基础与实验室效度。 它立在 orienting response 的心理物理学上(Sokolov 1963 谱系),元分析效应量 d=1.55(169 个条件、80 项实验室研究;模拟犯罪子集 d=2.09;最优 10 条件 d=3.12)(Ben-Shakhar & Elaad 2003,作者自存全文[一手逐字])。日本把 CIT 作为刑事侦查中唯一使用的测谎方法——「Japan is the only country where the polygraph with the concealed information test (CIT) is widely applied to criminal investigations」,约 100 名检查员年处理约 5,000 案(Matsuda et al. 2019, Front. Psychiatry[一手逐字])。
守真锚四:它是真实运转、被多国制度化的技术-制度复合体。 美国联邦政府约 1,100 名持证检测员、服务 30 个联邦机构,「conduct more than 110,000 screening, operational, and criminal specific examinations per year」(DCSA FY2026 预算论证,NCCA 节[一手逐字]);CBP 依 2010 年《反边境腐败法》对所有执法职位申请人法定强制测谎(P.L. 111-376[一手逐字]);英国把测谎写进性犯罪者、恐怖主义犯罪者与家暴犯罪者的假释条件(Offender Management Act 2007 s.28 现行合并文本[一手逐字])。
守真锚五:威慑与诱供效用被官方承认——但承认者自己把它和效度切开。 NAS 2003 执行摘要承认筛查可能有「some utility for such purposes… which are distinct from actual validity or accuracy」[一手逐字]——「让人招供」与「测得准」是两件事,委员会把这句话印在报告里。
守真锚六:批评者阵营的最强代表也承认个案用途。 Lykken——测谎批评谱系里地位最高的人——的立场是:「the polygraph is not a ‘lie detector’ but simply a recorder of physiological responses to verbal stimuli」,同时「There is empirical evidence to support its use in the investigation of specific incidents」(FAS 2006 年纪念文转述并嵌引《A Tremor in the Blood》[一手逐字])。他反对的是筛查——称之为「a menace in American life」——不是仪器本身。
四、仪器账:这台机器到底测什么(信号跳·物理层)
4.1 三个通道,没有一个是「谎言通道」
NAS 2003 第 3 章对仪器测量的描述逐字:「The polygraph machine usually measures three or four responses. Relative blood pressure is measured by a blood pressure cuff positioned over the biceps. Electrodermal activity (a measure of the activity of the eccrine sweat glands) is measured by electrodes placed on two fingers or the palm of the hand … The rate and depth of respiration are measured by pneumographs positioned around the chest and abdomen.」(NAS 2003 第 3 章[一手逐字])
三个通道的生理学性质:皮肤电活动(EDA)由交感神经支配的小汗腺驱动(附录 D 逐字:「innervated by the sympathetic branch of the autonomic nervous system, but the postganglionic neurotransmitter is acetylcholine rather than norepinephrine」[一手逐字])——它是实验室欺骗检测研究中最敏感的单一指标(「The most sensitive measure in laboratory studies of the detection of deception has been electrodermal activity」[一手逐字])。但它同时也是最没有特异性的指标:「stimuli that elicit the responses are so numerous as to make it difficult to isolate its specific psychological antecedent」[一手逐字]——能激发皮肤电的心理前因多到无法分离;而且呼吸是自主可控的,「a sharp sniff can reliably produce an electrodermal response」(猛地吸一下鼻子就能可靠地造出一个皮肤电反应)[一手逐字]。
这就是信号跳的物理基底:仪器测的是交感唤醒,而交感唤醒没有语义。恐惧、愤怒、被冤枉的羞辱感、对测试本身的焦虑、甚至一次深呼吸,都写在同一条曲线上。
4.2 从业者的自认与两份官方文件相隔二十年的同一句否定
仪器不测欺骗,这不是批评者的指控,而是从业共同体的正式立场。NAS 2003 第 1 章逐字:「The physiological phenomena that the instrument measures and that the chart preserves are believed by polygraph practitioners to reveal deception. Practitioners do not claim that the instrument measures deception directly. Rather, it is said to measure physiological responses that are believed to be stronger during acts of deception than at other times.」[一手逐字]
同章还有一句更狠的结构句:「A polygraph test and its result are a joint product of an interview or interrogation technique and a psychophysiological measurement or testing technique. It is misleading to characterize the examination as purely a physiological measurement technique.」[一手逐字]——测谎是一次审讯话术与一份生理测量的联合产品,把它描述成纯测量技术是误导。
欺骗无特异生理反应,官方文件说了两遍,隔了整整二十年:
- OTA 1983(美国国会技术评估办公室,OTA-TM-H-15):「The polygraph instrument, it should be noted, is not a ‘lie detector’ per se; i.e., it does not indicate directly whether a subject is being deceptive or truthful. There is no known physiological response that is unique to deception」(FAS 全文镜像[一手逐字])。
- NAS 2003(国家研究理事会):「There is no unique physiological response that indicates deception (Lykken, 1998). If deceivers in fact have stronger differential responses to relevant questions, it does not necessarily follow that an examinee who shows this response pattern was lying」(第 3 章[一手逐字])——后半句是逻辑层的点睛:即使说谎者平均而言反应更强,也不能反推出「反应强的人正在说谎」。这正是信号跳的形式化否定。
4.3 这个批评有多老:1917 年 12 月 14 日
它不是后现代解构,它跟这项技术同岁。1917 年,Marston 正在向国家研究理事会(NRC)兜售他的收缩压欺骗测试;同年 12 月 14 日,心理学家 John F. Shepard 在致委员会主席 Yerkes 的信里写下(NAS 2003 第 3 章逐字转引):
“The net result has been, I think to show that organic changes are an index of activity, of ‘something doing,’ but not of any particular kind of activity . . . but the same results would be caused by so many different circumstances, anything demanding equal activity (intelligence or emotional) that it would be impossible to divide any individual case.”
(机体变化是「有事情在发生」的指标,但不是任何特定种类事情的指标;同样的结果可由太多不同的情境产生,以致对任何个案都无法切分。)[一手逐字,NAS 转引自 NRC 档案未刊信]
1917 年的「something doing」、1983 年的「no known physiological response that is unique to deception」、2003 年的「no unique physiological response」——三句话相隔 86 年,说的是同一件事。信号层的天花板在第一天就写在那里,此后一百年没有被抬高过一毫米。
4.4 仪器史:三个发明者,一台机器
现代多道仪的谱系(NAS 2003 附录 E 逐字):「The polygraph literature variously attributes the origins of the modern polygraph machine to Benussi (1914) or to Larson, who constructed the prototype of the multi-channeled polygraph in 1921 … and to Keeler (1933).」[一手逐字]
- Marston(1915–1921):哈佛研究生期间做收缩压欺骗测试,1917 年向 NRC 提案(NRC「never officially hired Marston nor sponsored his work」[一手逐字])。他是「用心生理记录测欺骗」这个想法在法律与实验室场景的源头,但从未造出多通道仪器。
- Larson(1921):加州大学伯克利分校医学生,为警察局长 August Vollmer 组装了美国第一台多道仪。伯克利 Bancroft 图书馆档案指南逐字:「In 1921 John A. Larson, a young medical student at the University of California Berkeley, assembled the first American polygraph lie detector for Berkely Chief of Police, August Vollmer.」(OAC 典藏指南[一手逐字;原文拼写「Berkely」照录])。他给机器起的名字是 cardio-pneumo-psychogram(心肺心理描记器),「and later simply a polygraph, a nod to the multiple physical signals recorded by the stylus」(Smithsonian Magazine 2024[一手逐字])。
- Keeler(1933 起):定型多道仪并商业化——NAS 附录 E 逐字:「Keeler, by contrast, patented the hardware for his polygraph machine, controlled who could buy the machines, and marketed his approach to business and government; he did not systematically subject it to peer review.」[一手逐字]。他还在 1946 年把测谎引入橡树岭核设施,开创了安全筛查这个市场(同附录,引 Alder 1998)[一手逐字]。
值得注意的是 NAS 附录 E 顺手记下的一对行为对照:「Larson chose an ‘open science’ strategy … Throughout his career, he publicly expressed doubts about the suitability of polygraph tests as evidence in the courts. Keeler, by contrast, patented the hardware …」[一手逐字]——走科学发表路线的那位终生公开怀疑测谎出庭;走专利营销路线的那位把它卖给了政府和企业。仪器史的第一页就写好了此后一百年的分工:造它的人怀疑它,卖它的人不送审。
4.5 仪器层小结
仪器层裁决:多道仪是一台忠实的唤醒记录仪,它的三个通道没有一个对「欺骗」特异;这一点从业者自认、国会办公室 1983 年写明、国家科学院 2003 年重写、心理学家 1917 年预言。「测谎仪」这个名字里「谎」字所指向的那个通道,在物理上不存在。
五、提问账:三种问法,三种理论(信号跳·逻辑层)
仪器只给曲线,提问技术才给曲线以「意义」。测谎一百年先后有三种主格式,它们的分歧不在话术细节,在于它们对「曲线为什么说谎时会更陡」给出了三种互不相容的回答。
5.1 RIT:相关-无关测试(Marston 的恐惧说)
最早、曾长期主导的格式(NAS 附录 A:「The relevant-irrelevant test format was the first widely used polygraph testing format and was long the dominant format」[一手逐字])。理论逐字:「a guilty person, who is deceptive only to the relevant questions, will react more to those questions; in contrast, an innocent person, who is truthful about all questions, will not respond differentially」(NAS 第 3 章[一手逐字])。心理学根基是 Marston 1917 年定的调:「Marston (1917) described the underlying psychological state as fear; other writers have conceived it as arousal or excitement」[一手逐字]。
它的死刑判决由从业者自己签发:「Polygraph researchers generally consider the test outmoded. For example, Raskin and Honts (2002:5) conclude that the relevant-irrelevant test ‘does not satisfy the basic requirements of a psychophysiological test and should not be used.’」(NAS 附录 A[一手逐字])——无关问题与相关问题之间的反应差对「无辜但紧张的人」毫无保护,RIT 因此被本行业废弃。
5.2 CQT:对照问题测试(让无辜者对别的问题更紧张)
当今北美实务主流。它的核心机关是对照问题:「The comparison questions are specially formulated during a pretest interview with the intent to make an innocent examinee very concerned about them and either lie with high likelihood (a probable lie comparison question) or lie under instruction (a directed lie comparison question)」(NAS 第 3 章[一手逐字])。
判定逻辑逐字:「The theory is that the innocent person will show equal or less physiological responsiveness to relevant than comparison questions and that the guilty person will show greater responsiveness to relevant than comparison.」[一手逐字]——无辜者被设计为对对照问题(如「你这辈子有没有撒过谎」)反应更强,有罪者被预期对相关问题反应更强。
这个机关的承重件是话术,不是仪器。OTA 1983 逐字记下从业者 Raskin 的自认:「Control questions are intentionally vague and extremely difficult to answer truthfully with an unqualified ‘No’.」对照问题被故意设计得含糊、且几乎无法问心无愧地回答一个不带限定词的「没有」。OTA 接着写:「The polygraph examiner does not tell the subject that there is a distinction between the two types of questions (control and relevant). … the examiner wants the subject to experience considerable doubt about his or her truthfulness or even to be intentionally deceptive.」(OTA 1983[一手逐字])——考官不会告诉受测者两类问题有别;测试的有效性依赖于受测者被骗。CQT 因此是一种「以欺骗测欺骗」的装置:它的逻辑前提(无辜者会对对照问题更焦虑)必须由考官在测前访谈中亲手制造出来。NAS 据此写下一个结构性后果:「Since the subject’s psychological set is so crucial when control questions are used, differential responding … depends on the nature of the interaction between examiner and subject」([OTA 1983 同页逐字])——测试的效度取决于考官与受测者之间的互动质量,这使 CQT 更像一门手艺而不是一个仪器读数。
CQT 的理论谱系(冲突理论、心理定势理论、惩罚威胁理论)有一个共同软肋,NAS 第 3 章逐字:「they also predict that truthful examinees, under certain conditions, will show physiological response patterns similar to those expected from deceptive examinees … the polygraph test may not be specific to deception because other psychological states … mimic the physiological signs of deception.」[一手逐字]——CQT 自己的理论预言了假阳性:一个害怕被误判的无辜者,在理论上就会呈现出欺骗者的曲线。
5.3 CIT/GKT:隐蔽信息测试(识别,而不是欺骗)
Lykken 1959 年提出的另一条路:不问「你是不是撒谎了」,问「这件事的细节你是不是知道」。NAS 第 3 章逐字:「Lykken (1959, 1998) devised the guilty knowledge test (called here the concealed information test), based in part on orienting theory.」[一手逐字]理论基础是 Sokolov 的朝向反射(orienting response):「An orienting response occurs in response to a novel or personally significant stimulus」[一手逐字]——人对「对自己有意义的刺激」产生特异生理组合反应,与是否说谎无关。Ben-Shakhar 与 Furedy 称之为「a cognitive approach, because it emphasizes the fact that an individual knows something, rather than the individual’s emotions, concerns, fears, conditioned responses, or deception」(Ben-Shakhar & Elaad 2002 手册章,作者自存[一手逐字])。
它自带一个干净的基数逻辑(NAS 第 3 章逐字):无辜者对五选一题目中正确答案反应最强的概率是 1/5;若连续五题都对正确项反应最强,无隐蔽信息时该结果纯属巧合的概率是「1 in 5^5 (0.00032)」[一手逐字]——CIT 的推断是一个可计算的组合概率,不依赖考官话术。
理论分歧由 NAS 白纸黑字定性:「The claim that orienting theory provides justification for the comparison question technique of polygraph testing is radically at odds with the practices of polygraph examiners using that technique. … Thus, we do not take very seriously the argument that the TES or other polygraph examination procedures based on the comparison question technique can be justified in terms of orienting theory.」(第 3 章[一手逐字])——用朝向理论给 CQT 背书,与 CQT 的实际操作「根本对立」。CQT 测的是对谎言后果的恐惧,CIT 测的是对秘密信息的识别——这是两种不同的东西,共用同一台机器,却常被传播成同一件事。
5.4 委员会对理论层的总裁定:atheoretical
NAS 2003 第 3 章结论节逐字:
“The bulk of polygraph research can accurately be characterized as atheoretical. The field includes little or no research on a variety of variables and mechanisms that link deception or other phenomena to the physiological responses measured in polygraph tests.”
“Research on the polygraph has not progressed over time in the manner of a typical scientific field. Polygraph research has failed to build and refine its theoretical base, has proceeded in relative isolation from related fields of basic science … As a consequence, the field has not accumulated knowledge over time or strengthened its scientific underpinnings in any significant manner.”
“There has been no serious effort in the U.S. government to develop the scientific base for the psychophysiological detection of deception by the polygraph or any other technique, even though criticisms of the polygraph’s scientific foundation have been raised prominently for decades. The reason for this failure is primarily structural.”
(均出自NAS 第 3 章[一手逐字])
三句连读是一副完整的问责结构:研究无理论(atheoretical)→ 领域不积累(failed to build)→ 政府不投入(no serious effort)→ 原因是结构性的(primarily structural)。最重的不是「测谎不准」,而是这个领域一百年没有按科学的方式积累过知识。
行业一侧 2015 年的自我修正为此背书。APA 官方刊物上从业者 Nelson 的综述逐字承认:CQT 长期使用的「psychological set」(心理定势)模型「now regarded as an inadequate model for both polygraph responses and stress responses in general」,且「requires the assumption that polygraph sensors can identify different types of emotions, though the literature does not support this notion」(Nelson 2015, Polygraph 44(1)[一手逐字])——行业自己的理论旗手在 2015 年宣布:那个用了六十年的理论模型不够用,而且它要求传感器能区分情绪种类,文献不支持这个假设。
5.5 提问层小结
提问层裁决:RIT 被本行业废弃;CQT 的承重件是考官制造的「心理定势」——一种依赖欺骗受测者的话术,其理论被官方委员会判为 atheoretical,其操作模型被行业自己在 2015 年宣布 inadequate;CIT 有真实理论基础与可计算基数逻辑,但它的效度证据与「美国实务主流是 CQT」这个事实之间存在一条五十年的缝(第十三章展开)。信号跳在逻辑层的形态因此清楚了:机器后面没有一个「谎言理论」,只有三套互相打架的「为什么曲线会动」的假说,而被选中的那套(CQT)恰好是最依赖审讯话术的那套。
六、效度账:个案调查里它有多好(信号跳与分母跳的接缝)
先给结论形状:个案场景下测谎显著好于随机,但好于随机是一个远低于司法定论与筛查可用门槛的弱命题;而且你能看到的所有效度数字,都被三套系统性偏差往高估的方向推。这一层是信号跳与分母跳的接缝:个案效度是真有的,但它的数字口径撑不起它被用于的那些场景。
6.1 NAS 2003 的元分析:A 中位 0.86,且「最可能高估」
委员会从文献中筛出 57 项量化研究——执行摘要逐字:「Of the 57 studies the committee used to quantify the accuracy of polygraph testing, all involved specific incidents, typically mock crimes」(执行摘要[一手逐字])。这句「全部是个案研究」是后面整个分母账的地基:筛查场景根本没有可量化的研究基础。
效度用 ROC 曲线下面积 A 表示(「Its possible range is from 0.5 at the ‘chance’ diagonal to 1.0 for perfection」,第 2 章[一手逐字]):
- 实验室 52 个数据集:「the interquartile range of values of A reported for these data sets is from 0.81 to 0.91. The median accuracy index in these data sets is 0.86」(第 5 章[一手逐字])。
- 现场研究只有 7 项过质量关:「The seven datasets include between 25 and 122 polygraph tests, with a median of 100 and a total of 582 tests. … The accuracy index values (A) range from 0.711 to 0.999, with a median value of 0.89」——但这个 0.89 与实验室的 0.86「statistically indistinguishable」[一手逐字]。
- 研究间散布极大:假阳性率固定约 10% 时,敏感度「ranges from 43 to 100 percent」[一手逐字]——同一标称假阳性率下,抓真凶的能力从四成出头到满分不等。这就是「0.86」这个中位数所掩盖的真实世界。
然后是委员会亲手写下的降格:「we believe that the range of accuracy indexes (A) … with a midrange between 0.81 and 0.91, most likely over-states true polygraph accuracy in field settings involving specific-incident investigations. We remind the reader that these values of the accuracy index do not translate to percent correct」(第 5 章[一手逐字])。并给出上界判定:「reliance on polygraph testing to perform in practical applications at a level at or above A = 0.90 is not warranted on the basis of either scientific theory or empirical data. Many committee members would place this upper bound considerably lower.」[一手逐字]——注意最后这句:连 0.90 这个上界,都有多名委员认为定高了。
换算成正确率口径:A=0.9 在平衡阈值下对应 81.8% 正确,A=0.7 对应 66.0%(第 2 章逐字)[一手逐字]。执行摘要再补一刀:「Estimates of accuracy from these 57 studies are almost certainly higher than actual polygraph accuracy of specific-incident testing in the field.」[一手逐字]
6.2 为什么看到的数字系统性偏高:confession 偏差
现场研究用「认罪」来确认真值,这个做法本身就制造高估。NAS 第 4 章逐字:「Such research necessarily omits cases in which there was no confession. This procedure probably yields an upward bias in the estimates of polygraph accuracy because the relationship between polygraph results and guilt is likely to be stronger in cases that led to confessions」(第 4 章[一手逐字])。
委员会给了一个定量演示:真实敏感度 67% 的测试,经认罪确认样本筛选后会被读出 86%;真实 80% 的会被读出 92%(第 4 章定量例[一手逐字])。Iacono 把这个选择机制写得更直白:「the only cases selected for study would be those where the original examiner was both correct and obtained a verifying confession」(Iacono 2001,JFPP 全文镜像[一手逐字])——只有考官判对且拿到认罪的案子才进得了「效度研究」的样本。效度表上的每一个百分点,都可能只是「认罪倾向与曲线形态相关」的倒影。
实证定锚:Patrick & Iacono 1991 对加拿大皇家骑警(RCMP)现场数据的研究(经 Iacono 2001 逐字转述):「the accuracy of the CQT with confession verified innocent suspects was only 57%. Second, it was not possible to estimate accuracy for guilty subjects because once a suspect failed a polygraph test, even with no confession, the police generally stopped investigating the case」[一手逐字]——无辜者准确率 57%(略高于抛硬币),而有罪者准确率根本无法估计:一旦嫌疑人测谎失败,警方就停止侦查,真值永远沉没。
6.3 OTA 1983:更早二十年的官方评估,同一形状
国会技术评估办公室 1983 年的评估(应众议院政府运作委员会主席 Jack Brooks 与资深委员 Frank Horton 委托,前言逐字[一手逐字,扫描 OCR 件]):
- 10 项个案现场研究汇总:有罪检出 70.6–98.6%(均值 86.3%);无辜者正确检出 12.5–94.1%(均值 76%);假阳性率 0–75%(均值 19.1%);假阴性率 0–29.4%(均值 10.2%)(第 7 章 p.97,逐字,OCR 件)。
- 结论句:「The preponderance of research evidence does indicate that, when the control question technique is used in specific-incident criminal investigations, the polygraph detects deception at a rate better than chance, but with error rates that could be considered significant.」(p.97–98)[一手逐字]
- 总限定:「OTA concluded that there is at present only limited scientific evidence for establishing the validity of polygraph testing.」[一手逐字]
- 一个值得存档的彩蛋:OTA 引述了 FBI 自己的内部结论——「even the Federal Bureau of Investigation (FBI) has concluded that, ‘to date, no methodologically adequate study of control question techniques has been reported…’」(p.98)[一手逐字]。1983 年,连 FBI 都承认 CQT 没有方法学合格的研究。
无辜者正确检出率均值 76%、区间下限 12.5%——把这两个数字并排放着看,就能看到「个案效度」的真实形状:对有罪者它相当灵敏(86.3%),对无辜者它的保护力区间下沿可以低到 1/8。机器的两类错误从来不对称,而被假阳性击中的代价由无辜者承担。
6.4 科学共同体的意见调查:约三分之一认可
Iacono & Lykken 1997(J. Applied Psychology 82(3):426–433)调查了两个精英样本:心理生理学研究学会(SPR)会员(回收率 91%)与美国心理学会第一分会会士(回收率 74%)。原文付费墙未破(附录 B 登记),数字经同一第一作者 2001 年的同行评审复述与联邦司法中心《科学证据参考手册》第三版双源交叉:「Only about a third of those surveyed believed the CQT was scientifically sound, and only about a quarter thought polygraph evidence should be admissible in court. Few opined that the CQT was a standardized test (20%) or objective (10%).」(Iacono 2001[一手逐字];FJC 手册第三版 第 367 页起含页码指引[一手逐字])
FJC 手册同时记下了调查口径之争:早期调查中「约 60%」受访者认为测谎是「a useful diagnostic tool when considered with other available information」,但 Iacono 指出该问题「dealt neither with the CQT nor its use in court」[一手逐字]——「和其他信息一起看时有点用」与「科学上成立」是两个问题,传播链只取了前一个。这是本篇最小的那个跳变样本:同一份问卷,读「useful」者得 60%,读「scientifically sound」者得三分之一。
6.5 行业元分析:.890 与它自己的一句拆台
行业协会(American Polygraph Association)2011 年发表了自己的元分析(Polygraph 40(4),118 页,官方 PDF[一手逐字]):
- 事件特定单议题诊断:聚合决策准确率 .890(.829–.951),不确定率 .110;多议题 .850(.773–.926);全体「validated」技术 .869(.798–.940)。
- 结论句:「Data at the present time are sufficient to support the polygraph as highly accurate, but insufficient to support an assertion that PDD testing can provide perfect or near-perfect accuracy.」
同一页里还有一句值得单独装裱的自我拆台:「This illustrates that the APA categorical distinctions are arbitrary, not empirically founded, and scientifically meaningless. These data are insufficient to support the notion that any PDD technique is superior to another…」[一手逐字]——行业元分析自己承认:技术之间的分类区别是武断的、无实证基础、科学上无意义。
两套数字的对峙留在这里就够了:APA 2011 聚合 .890 vs NAS 2003「A≥0.90 的现场表现 not warranted」。从业阵营内部最新的元分析(Honts, Thurber & Handler 2021,138 数据集)含不确定结局的效应量只有 0.69(OpenAlex 逐字摘要;全文付费墙未破,附录 B 登记[需亲核])。
6.6 效度层小结
个案层裁决:好于随机(A 中位 0.86,「well above chance, though well below perfection」)——这句官方结论同时是两个阵营的弹药库。支持方取前半句,批评方取后半句与那三个限定词(研究文献人群、未受反措施训练、个案场景)。而委员会自己写的三行降格(「most likely over-states」「almost certainly higher」「Many committee members would place this upper bound considerably lower」)告诉你:就连 0.86 这个数,也是含着系统性高估水分的数。个案效度是真的;它的口径撑不起它被用于的任何大规模场景——这就是分母跳的入口。
七、分母账:10,000 人里抓 10 个间谍的算术(本篇最重格)
个案效度再好,也只是「这个嫌疑人是不是在撒谎」的效度。国家安全筛查面对的是另一个问题:「这一万名雇员里有没有间谍」。两个问题之间隔着基数(base rate)——而基数是算术,不是观点。NAS 2003 委员会做了一件此后很少见到的事:它亲手把这笔算术算给你看了。
7.1 委员会的自算表:表 S-1 逐字还原
执行摘要印刷页第 5–6 页,表 S-1(NAS 2003 执行摘要[一手逐字])。标题本身逐字:「Expected Results of a Polygraph Test Procedure with an Accuracy Index of 0.90 in a Hypothetical Population of 10,000 Examinees That Includes 10 Spies」——注意它用的是 A=0.90,一个比实测中位(0.86)还高的乐观上界。
阈值设为「抓住 80% 的间谍」时:
| 测试结果 | 间谍 | 非间谍 | 合计 |
|---|---|---|---|
| 「不通过」 | 8 | 1,598 | 1,606 |
| 「通过」 | 2 | 8,392 | 8,394 |
| 合计 | 10 | 9,990 | 10,000 |
阈值改为「大幅压低假阳性」时:
| 测试结果 | 间谍 | 非间谍 | 合计 |
|---|---|---|---|
| 「不通过」 | 2 | 39 | 41 |
| 「通过」 | 8 | 9,951 | 9,959 |
| 合计 | 10 | 9,990 | 10,000 |
表格正文前后的说明(p.6,逐字):
“If the test were set sensitively enough to detect about 80 percent or more of deceivers, about 1,606 employees or more would be expected ‘fail’ the test; further investigation would be needed to separate the 8 spies from the 1,598 loyal employees caught in the screen. If the test were set to reduce the numbers of false alarms (loyal employees who ‘fail’ the test) to about 40 of 9,990, it would correctly classify over 99.5 percent of the examinees, but among the errors would be 8 of the 10 hypothetical spies, who could be expected to ‘pass’ the test and so would be free to cause damage.”
同页结论句(逐字):
“Given its level of accuracy, achieving a high probability of identifying individuals who pose major security risks in a population with a very low proportion of such individuals would require setting the test to be so sensitive that hundreds, or even thousands, of innocent individuals would be implicated for every major security violator correctly identified.”
以及执行摘要那条大写 CONCLUSION(逐字):
“CONCLUSION: Polygraph testing yields an unacceptable choice for DOE employee security screening between too many loyal employees falsely judged deceptive and too many major security threats left undetected. Its accuracy in distinguishing actual or potential security violators from innocent test takers is insufficient to justify reliance on its use in employee security screening in federal agencies.“
这就是分母跳的全部内容:同一台仪器、同一份效度,把基数从 1/2(刑事嫌疑人)换成 1/1,000(国家实验室雇员),「抓间谍」就变成「错标清白者」。阈值往左调,错标 1,598 人;往右调,放走 8 个间谍。没有第三个旋钮。
7.2 「99.5% 正确」的陷阱与 PPV 账
第 2 章的表 2-1(印刷页 p.48–49)把对照摆得更细:同一个 A=0.90 的测试、三档灵敏度、两个人群并排。表前引导句(p.48,逐字):
“In a population of 10,000 criminal suspects of whom 5,000 are expected to be guilty, the test will identify 4,800 examinees (on average) as deceptive, of whom 4,000 would actually be guilty. The same test, used to screen 10,000 government employees of whom 10 are expected to be spies, will identify an average of 1,606 as deceptive, of whom only 8 would actually be spies.”
然后是那句值得刻在每一张「准确率 99.5%」广告旁边的陷阱句(p.49 附近,逐字):
“The problems with percent correct as an index of accuracy are best seen in the situation shown in the right half of Table 2-1c, in which the test is correct in 9,959 of 10,000 cases (99.5 percent correct), but eight out the ten hypothetical spies ‘pass’ and are free to cause damage.”
「99.5% 正确」与「10 个间谍放走 8 个」是同一张表的同一行。任何用 percent correct 给筛查测谎背书的说法,都在这张表里现形。
第 7 章把这笔账推广成指标(NAS 2003 第 7 章,pp.179–185,[一手逐字]):
- 假阳性指数 FPI=每个真阳性对应的假阳性数。基数 1/1,000 时:「an individual who is judged deceptive on the test will in fact be nondeceptive more than 199 times out of 200, even if the test has A = 0.90, which is highly unlikely for the polygraph.」——被判「欺骗」的人,超过 199/200 的概率是清白的。
- 刑事对照:「For a test with A = 0.80 and a sensitivity of 50 percent, the false positive index is 0.23 and the positive predictive value is 81 percent. That means that someone identified by this polygraph protocol as deceptive has an 81 percent chance of being so, instead of the 0.4 percent (1 in 250) chance of being so if the same test is used for screening a population with a base rate of 1 in 1,000.」——同一个测试,刑事场景下「阳性即 81% 真有事」,筛查场景下「阳性即 0.4% 真有事」。差 200 倍的不是仪器,是分母。
- 注 4(p.211)给出敏感性:基数 ≤1/1,000 时 FPI 至少为 208(A=0.90)/452(A=0.80)/634(A=0.70)/741(A=0.60);若 10,000 名雇员中有 10 个重大违规者且阈值设为抓住其中 8 个,则误标为欺骗的清白者至少为 1,664/3,616/5,072/5,928 人。
- 分母本身的数量级(pp.184–185,逐字):「The one major spy caught in the FBI is one among perhaps 100,000 agents who have been employed in the bureau’s history. The base rate of major security threats in the nation’s security agencies is almost certainly far less than 1 percent.」
7.3 「从个案外推到筛查不正当」:委员会自己把话说死
委员会没有留下任何「也许可以」的缝隙。三句直接句(均逐字):
“Because the studies of acceptable quality all focus on specific incidents, generalization from them to uses for screening is not justified. Because actual screening applications involve considerably more ambiguity for the examinee and in determining truth than arises in specific-incident studies, polygraph accuracy for screening purposes is almost certainly lower than what can be achieved by specific-incident polygraph tests in the field.”(执行摘要 p.4)
“Of the 57 studies the committee used to quantify the accuracy of polygraph testing, all involved specific incidents, typically mock crimes…”(执行摘要 p.3)
“To summarize, the performance of the polygraph is sharply different in screening and in event-specific investigation contexts. Anyone who believes the polygraph ‘works’ adequately in a criminal investigation context should not presume without further careful analysis that this justifies its use for security screening.“(第 7 章 p.185)
第三句值得单读:就算你在个案场景相信它「够用」,委员会明令——不许把这份相信带进筛查场景。
7.4 一次罕见的落地:DOE 2006 终规则
科学报告的结论通常进了新闻通稿就结束。这一份进了联邦法规——而且是国会用法条押着它进去的。NDAA FY2002 §3152(Pub. L. 107-107,编入 42 U.S.C. 7383h-1)指令能源部制定新的反情报测谎规则,并且(71 FR 57386,2006-09-29 终规则[一手逐字])「take into account the results of the Polygraph Review」——法条里把「the Polygraph Review」定义为国家科学院这个委员会的审查。一份科学评估被国会点名写进立法指令,这本身就是罕见样本。
DOE 在终规则里的政策句(逐字):
“Consistent with the practices of the Intelligence Community, and the NAS Report, DOE has decided to alter the role of polygraph testing as a required element of the counterintelligence evaluation program by eliminating such testing for general screening of applicants for employment and incumbent employees without specific cause.”
保留五种情形(外国联系疑点、外派要求、他机构请求、随机抽测、个案调查),并明言:「That proposal also grew out of the NAS Report, which noted that this kind of use of the polygraph is the one for which the existing scientific literature provides the strongest support.」(句中 this kind of use 指个案用途)[一手逐字]
这是报告落地的样子:法定指令 → 规则修订 → 「无具体事由的一般筛查」被明文取消。 1999 年 DOE 按 NDAA FY2000 §3154 建立的 10 CFR 709 八类岗位强制测谎(64 FR 70962[一手逐字]),七年后被同一部门按 NAS 报告收缩。
7.5 但总量不减:从 11,546 到 110,000+
DOE 收缩了,联邦机器没有。规模账按年代排开(全部一手):
- DoD FY1999(国防部致国会年报,全文重刊于 Polygraph 29(3),APA 期刊 PDF[一手逐字]):全年 11,546 次测谎(不含 NSA 与 NRO),其中反情报筛查(CSP)8,289 次——而国会给 CSP 的法定上限是每年 5,000 次(FY1988–90 为 10,000;FY1991 起 5,000,P.L. 100-180 §1121),超出的 5,045 次走豁免通道。产出端更值得看:8,088 人「无显著反应」,201 人「显著反应或存疑或提供实质信息」——其中 187 人获有利裁定,8 人待查,5 人待定,仅 1 人被采取不利行动;逐字一句单独存档:「There was one individual whose polygraph examination result was evaluated as significant physiological response (deceptive) and who made no admissions to the relevant issues.」——全年近万次筛查,唯一一个「测了、欺骗、不招」的人,案卷里就这一个。
- FBI FY2002–2005(DOJ OIG 2006 报告[一手逐字,Wayback 快照]):四年 38,017 次,年均增长约 45%;FY2005 单年超 11,000 次;79% 是入职前与人员安全筛查;强制定期/随机测谎的覆盖人群从 2001 年的 550 人膨胀到 2005 年的 18,384 人。
- 现役总量(DCSA 预算论证,NCCA 节):约 1,100 名联邦训练持证检测员、服务 30 个联邦伙伴机构,「conduct more than 110,000 screening, operational, and criminal specific examinations per year」(FY2026 预算件[一手逐字];FY2024 版同文件为「more than 103,000」,FY2024 预算件[一手逐字])。
- CBP:2010 年《反边境腐败法》(P.L. 111-376)把测谎从「机构裁量」升级为「法定强制」——标题逐字即「To require U.S. Customs and Border Protection to administer polygraph examinations to all applicants for law enforcement positions」(govinfo[一手逐字])。落地后的通过率:FY2015→FY2017,边境巡逻队员 28%→26%,CBP 官员 32%→25%;GAO 证词称测谎「has consistently had the lowest pass rate of any step in its hiring process」(GAO-19-419T[一手逐字])。约七成申请者倒在测谎环节——在一个委员会认定筛查场景假阳性远多于真阳性的测试上。
- 地方层:BJS《Local Police Departments, 2007》(NCJ 231174)附录表 4:26% 的地方警察部门要求测谎;服务人口 100 万以上的辖区为 77%,25–50 万辖区为 83%(报告 PDF[一手逐字])——按部门数算是少数,按覆盖警员数算是多数,两个分母两种读法。
时间轴合起来读:2003 年委员会判「筛查不可用」→ 2006 年 DOE 依法收缩 → 同一时期 FBI 筛查量翻倍、强制人群膨胀 33 倍 → 2010 年国会给 CBP 立法强制全员测谎 → 2020 年代联邦年测谎量突破 11 万次。报告的算术赢了 DOE 的规章,输掉了整个制度的总量。
7.6 EPPA 的结构不对称:被禁的是效度最强的场景
1988 年《雇员测谎保护法》(EPPA,29 U.S.C. ch.22)常被引为「美国禁止测谎」的证据。法条原文说的完全是另一回事(govinfo 法典[一手逐字]):
- §2002 禁止的是私人雇主:「it shall be unlawful for any employer engaged in or affecting commerce … to require, request, suggest, or cause any employee or prospective employee to take or submit to any lie detector test」——连「建议」员工去测都在禁止之列。
- §2006(a):「This chapter shall not apply with respect to the United States Government, any State or local government, or any political subdivision of a State or local government.」——政府整体豁免于自己签署的这部法律。
- §2006(b) 再点名四家:NSA、DIA、NGA、CIA 的雇员、合同方与申请人(以及接触 top secret/SAP 信息的联邦合同专家),在反情报职能下可施测。
把 EPPA 与 NAS 效度账叠在一起,不对称显形:个案刑事侦查(效度证据最强的场景)中,私人雇主——比如怀疑员工偷钱的店主——被法律禁止用它;大规模雇员筛查(效度证据最弱、委员会明判「insufficient to justify reliance」的场景)中,联邦政府不仅豁免,还对 CBP 申请人法定强制。效度最强的用法被禁给私人,效度最弱的用法被保留给国家。 立法的分界线从来不是效度,是谁在用。
7.7 分母层小结
分母层裁决:筛查不可用,不是批评者的修辞,是委员会的自算表、结论句与法定收缩规则的三重合议。表 S-1 的每一格都是算术:A=0.90(乐观上界)+基数 1/1,000 → 抓 8 个间谍要搭进 1,598 个清白者。而制度的回应是双轨的——DOE 按报告收缩(罕见的落地样本),联邦总量却从年 1.2 万次膨胀到 11 万次以上。分母跳在文件里被判死刑,在文件外照常运转。
八、间谍账:测谎没有抓到的那几个人
算术说「低基数下错标远多于真捕」,档案给出了实证版:美国测谎史上三宗最著名的间谍案,仪器的表现各不相同——共同点是,没有一个间谍是被测谎抓出来的。
8.1 Aldrich Ames:三次测谎,三次穿过
参议院特别情报委员会(SSCI)1994 年的 Ames 案评估报告(FAS 全文镜像[一手逐字])记录了三次:
- 1962 年 3 月 23 日入职测谎:「Ames passed the polygraph」。彼时他尚未为苏联工作——这次通过不是问题。
- 1986 年 5 月 2 日(赴罗马任职前,距他开始为 KGB 工作约一年):「Ames gave consistently deceptive responses to issues related to whether he had been ‘pitched’ … by a foreign intelligence service.」——仪器抓到了。然后呢?「This time, the CIA examiner deemed Ames truthful and concluded the examination… The examiner’s supervisors concurred with the assessment that Ames was non-deceptive.」——考官把持续欺骗反应解释掉了,上级附议。FBI 在 1993 年 6 月复核当年图谱后的结论逐字:「the deception indicated in Ames’s response to the pitch issue in 1986 was never resolved, even though the CIA examiner passed Ames on this exam. Also in the opinion of the FBI, significant deceptive responses by Ames were detectable to questions dealing with unauthorized disclosure of classified material.」——信号在纸上,判读把它放了过去。
- 1991 年 4 月 12 日:「Ames also showed no reaction when he was asked whether he was working for a foreign intelligence service.」——对「你是否在为外国情报机构工作」这个直球问题,这名当时已间谍六年的 CIA 军官毫无反应;只对「是否隐瞒与外国人的接触」显示欺骗。四天后换一名检测员复测:「Ames answered all questions without any deceptive response… The polygrapher concluded the test, dismissed Ames and wrote the polygraph report indicating Ames had passed the test.」首名检测员的存档批注是这个体系罕见的一句诚实:「I don’t think he is a spy, but I am not 100% convinced because of the money situation.」
委员会对体系后果的描述(逐字):审查专案组把 1991 年 4 月的通过结论当作给 Ames 的「a clean bill of health」(健康证明);并写下机制诊断:「All too often an officer who has been through training, gone through the polygraph examination, and had an overseas assignment, is accepted as a ‘member of the club‘, whose fitness for assignments, promotions, and continued service becomes immune from challenge.」——通过测谎成了俱乐部会员卡,一旦持卡,质疑免疫。
Ames 最终如何落网?不是测谎,是 1985 年起苏联线人接连暴露引发的反情报调查、他的奢靡消费与银行存款审计。仪器在他身上留下的唯一痕迹,是两次被考官放行的欺骗反应和一次毫无反应。
8.2 Robert Hanssen:二十五年,从未被测
Hanssen 案是另一种失败形态:不是「测了没测出来」,是「根本没测」。DOJ 监察长 2003 年审查报告解密执行摘要(32 页,Wayback 快照[一手逐字]):
“Hanssen also encountered few security checks at the FBI. He was never asked to submit to a polygraph examination or to complete a financial disclosure form, and he received only one background reinvestigation during his 25-year FBI career.”
更刺目的是 1994 年的内部史(同一报告 pp.21–22,逐字):
“After Ames’s 1994 arrest, FBI National Security Division managers argued for an aperiodic, random polygraph program, but the FBI’s most senior management rejected that request, largely because of concerns regarding false positives.”
——Ames 案发后,FBI 自己的国安部门提出随机测谎,最高层以「顾虑假阳性」否决。这个否决的理由恰恰是对的(第七章的算术),而被否决的对象恰恰是 Hanssen(当时正在卧底俄罗斯)。假阳性顾虑保护了真间谍——这不是反讽修辞,是 OIG 白纸黑字的因果链。
Hanssen 2001 年被捕后,FBI 才三步建起强制测谎:2002 年 4 月把反情报测谎纳入五年背景复查、2003 年 8 月增加随机测谎、2005 年 9 月扩展至全员(DOJ OIG 2007 复查报告[一手逐字,FOIA 镜像扫描件])。
NAS 委员会顺手做了反事实推演(第 7 章 p.187,逐字):「if Robert Hanssen had taken such tests three times during 15 years of spying, the chances are that, even without attempting countermeasures, he would not have been detected before considerable damage had been done.」——就算当年测了三次,委员会的判断也是:大概率还是测不出来。
8.3 Montes 与 Fuchs:一句「passed a polygraph」与一次阴性登记
- Ana Montes(DIA 古巴分析师,2001 年被捕):FBI 官方案例页逐字:「During her years at DIA, security officials learned about her foreign policy views and were concerned about her access to sensitive information, but they had no reason to believe she was sharing secrets. And she had passed a polygraph.」(FBI 官网[一手逐字])。其律师 2002 年对 Miami Herald 称她「至少一次」接受测谎且通过(二手转录,[需亲核])。测谎次数与年份一手司法文书未取回(附录 B 登记),但「通过」这一事实由 FBI 官方页确认。
- Klaus Fuchs(1950 年认罪的核间谍):MI5 官方历史页载明检出链条——VENONA 信号情报破译+安全官 William Skardon 的劝供面谈(「pressure would be put on Fuchs to induce him to confess… a month later he told Skardon that his conscience had compelled him to come clean」,Wayback 快照[一手逐字]);该页全文无「polygraph」出现。阴性登记:英国当时不以测谎为手段,缺席不等于绝对不存在,但官方叙事中测谎在检出链条里没有任何位置。
8.4 Lykken 的「盾牌」论
为什么通过测谎比不测更危险?Lykken 给出的机制解释(FAS 2006 年纪念文转述,一手逐字):通过测谎的人「become immune to commonsense suspicions」——测谎通过结论充当了一面盾牌,把本该持续的常理怀疑挡在门外。Ames 的「clean bill of health」与「member of the club」是这面盾牌的档案形态:仪器最危险的时刻不是它误报,是它放行——放行的纸面结论让制度停止追问。
8.5 间谍层小结
间谍层裁决:三宗大案三种形态——Ames(测了,信号在纸上,被判读放行)、Hanssen(二十五年没测,高层以对的理由否决了对的时机)、Montes(通过了,FBI 官方页把「passed a polygraph」写进案例史)。没有一例由测谎检出;NAS 委员会对「如果测了会怎样」的反事实答案也是「大概率还是检不出」。一个以「抓间谍」为首要正当性叙事的制度,在它的三宗最大考题上交了这样一份答卷——而同一时期,它每年给上万名清白雇员贴上「欺骗」标签。这就是分母跳在真实世界的价格表。
九、判例账:669 词如何统治美国证据法七十年
测谎仪对美国法律的最大贡献,不是抓到的任何人,而是一个判例。Frye v. United States——一份不足两页、全文 669 词、未引用任何先例的特区上诉意见——在此后七十年里成为美国科学证据可采性的总判据,直到 1993 年被最高法院亲手废除。而产生它的那台仪器,至今不被允许出庭作证。这是本篇最纯粹的判例跳样本:一个具体的、关于一台具体仪器的裁决,被读成了一条抽象的总规则。
9.1 Frye 全文:不足两页,零先例
1923 年 12 月 3 日,哥伦比亚特区巡回上诉法院,Van Orsdel 法官执笔,James Alphonso Frye 谋杀案二审(判决全文,CourtListener 经 Wayback[一手逐字])。全文在 Federal Reporter 上占 1013–1014 两页,不足两页;Lepore(2015)记其长度为「In a 669-word opinion」[一手逐字];判决未引用任何先例——Daubert 判决后来称它为「a short and citation free 1923 decision」。
案情只有一句(逐字):「A single assignment of error is presented for our consideration. In the course of the trial counsel for defendant offered an expert witness to testify to the result of a deception test made upon defendant. The test is described as the systolic blood pressure deception test.」——辩方要传一名专家就「收缩压欺骗测试」的结果作证,被初审拒绝。
然后是那段统治美国证据法七十年的规则陈述(全文仅此一段,逐字):
“Just when a scientific principle or discovery crosses the line between the experimental and demonstrable stages is difficult to define. Somewhere in this twilight zone the evidential force of the principle must be recognized, and while courts will go a long way in admitting expert testimony deduced from a well-recognized scientific principle or discovery, the thing from which the deduction is made must be sufficiently established to have gained general acceptance in the particular field in which it belongs.”
「普遍接受」(general acceptance)标准就此诞生——在一个关于测谎仪的、不点任何专家名字的、零引用的判决里。holding 逐字:
“We think the systolic blood pressure deception test has not yet gained such standing and scientific recognition among physiological and psychological authorities as would justify the courts in admitting expert testimony deduced from the discovery, development, and experiments thus far made.”
“The judgment is affirmed.”
测谎被挡在门外,Frye 维持二级谋杀罪成。判决甚至没说那名专家是谁(见 11.3)。
9.2 沉睡五十年:引用计量
判例跳的第一层不是「被引用太多」,是「先被遗忘五十年,再被当成自古以来的标准」。Faigman、Porter & Saks(1994,Cardozo L. Rev.,开放获取 PDF[一手逐字])的计量是这条曲线的官方账:
正文(1808 页,逐字):
“So unimportant was the Frye corollary that it went unnoticed for decades. No contemporary law review articles were written about it, commentators ignored it, and other courts did not cite it. In fact, Frye only became ‘trendy‘ in the 1970s as an argument arose over the admissibility of scientific evidence, perhaps in anticipation of the new Federal Rules.”
脚注 25(逐字,沉睡的量化):
“Frye was not cited by a single other court, federal or state, for a decade. During the first quarter century after its publication, Frye was cited in eight federal cases and five state cases. During its second quarter century, it was cited fifty-four times in federal cases and twenty-nine times in state cases. By the 1980s, it was being cited as much each year as it had been in its first fifty years added together.“
十年零引用;前 25 年合计 13 次;第二个 25 年 83 次;1980 年代每年的引用量约等于前 50 年总和。同文还点破了复活的经济学:「The intellectual or professional marketplace was simply a proxy for the commercial marketplace.」——Frye 把「商业市场检验」平移成了「知识市场检验」。一条为测谎仪量身定做的规则,在沉睡半个世纪后,被联邦证据规则立法前的争论从档案里捞出来,追封为开国元勋。
9.3 Daubert 1993:亲手废除
1993 年,最高法院在 Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579(Cornell LII 全文[一手逐字])中终结 Frye 在联邦系统的统治。关键句逐字:
“The Frye test has its origin in a short and citation free 1923 decision concerning the admissibility of evidence derived from a systolic blood pressure deception test, a crude precursor to the polygraph machine. In what has become a famous (perhaps infamous) passage…”
“Frye made ‘general acceptance’ the exclusive test for admitting expert scientific testimony. That austere standard, absent from and incompatible with the Federal Rules of Evidence, should not be applied in federal trials.“
“That the Frye test was displaced by the Rules of Evidence does not mean, however, that the Rules themselves place no limits on the admissibility of purportedly scientific evidence.”
「displaced」——取代,不是补充。Daubert 换上的是审判法官守门制(可检验性、同行评议、已知或潜在错误率、标准维护,「general acceptance」降级为参考因素之一)。669 词的统治到此结束:从 1923 到 1993,七十年,其中前五十年它在睡觉。
9.4 Scheffer 1998:8-1,「陪审团才是测谎仪」
Daubert 没有为测谎开门。五年后,最高法院在 United States v. Scheffer, 523 U.S. 303 中以 8-1 维持军事证据规则 707 对测谎证据的 per se 排除(多数意见[一手逐字])。Thomas 多数意见:
“The contentions of respondent and the dissent notwithstanding, there is simply no consensus that polygraph evidence is reliable. To this day, the scientific community remains extremely polarized about the reliability of polygraph techniques.”
“Although the degree of reliability of polygraph evidence may depend upon a variety of identifiable factors, there is simply no way to know in a particular case whether a polygraph examiner’s conclusion is accurate, because certain doubts and uncertainties plague even the best polygraph exams.”
“Such jurisdictions may legitimately determine that the aura of infallibility attending polygraph evidence can lead jurors to abandon their duty to assess credibility and guilt.”
脚注 7 顺手给了州格局:「Most States maintain per se rules excluding polygraph evidence. … New Mexico is unique in making polygraph evidence generally admissible without the prior stipulation of the parties and without significant restriction.」[一手逐字]
以及那句将被此后每一份测谎判决书引用的引语(Thomas 引第九巡回 1973 年 Barnard 案,逐字):
“A fundamental premise of our criminal trial system is that ‘the jury is the lie detector.'”
唯一异议者 Stevens 的反驳同样值得存档(异议意见[一手逐字]):Rule 707 连「双方事先书面同意可采」都禁止,「even if the results of the polygraph test were more reliable than the results of the urinalysis, the weaker evidence is admissible and the stronger evidence is inadmissible」;对陪审团能力的担忧「reflects a distressing lack of confidence in the intelligence of the average American」。
9.5 判例层小结:讽刺格
判例层裁决:Frye 的一生是三段不对称——669 词的轻量出生、五十年的沉睡、被追封后七十年统治。Daubert 废掉了它在联邦的名位,但守门人逻辑(错误率、标准、同行评议)落到测谎身上,结论与 1923 年相同:Scheffer 8-1 维持排除,州层面大多数 per se 排除(第十章)。产生「科学证据总判例」的那台机器,自己从未达到那个判例要求的门槛。 判例跳的完整形状:一次具体裁决被读成总规则,总规则统治七十年,而总规则的第一案标的物至今在门外。
十、法域账:拼布、例外与国际
10.1 州拼布:一份判决内嵌的 50 州普查
州层格局的最好来源不是学术综述,而是一份判决书的内嵌普查。Lee v. Martinez, 2004-NMSC-027(新墨西哥州最高法院,FindLaw 全文[一手逐字];案号勘误:任务书预设 051,正确为 027,附录 A 登记)逐字清点:
“Twenty-seven (27) states and the District of Columbia apply a per se rule of exclusion of polygraph evidence for all purposes.”
“Seventeen (17) states admit polygraph evidence at trial only when its admission is stipulated to in advance by all parties.”
“Two (2) other states admit stipulated results but in limited circumstances.”
路易斯安那与密歇根允许免约定在审判后程序采纳;南卡罗来纳交由初审法官按 702/403 裁量;新墨西哥常态可采。另有四州(马萨诸塞、北卡、俄克拉何马、威斯康星)曾采纳多年后回流 per se 禁止。联邦层面:仅第四巡回与特区巡回 per se 排除,「Most federal appellate courts leave admission of polygraph evidence to the discretion of the trial courts, but generally such evidence is excluded on the basis of Daubert/Rule 702 or Rule 403 or both.」[一手逐字]
「stipulation」(双方事先书面同意)规则的奠基判例是 State v. Valdez, 91 Ariz. 274 (1962)(CourtListener 经 Wayback[一手逐字]),四条件:①检察官、被告与辩护人三方签署书面约定;②审判法官保留裁量否决权;③对方有权就检查员资质、施测条件、技术局限与误差可能交叉询问;④法官须指示陪审团「the examiner’s testimony does not tend to prove or disprove any element of the crime … at most tends only to indicate that at the time of the examination defendant was not telling the truth」。注意 Valdez 存档的 1962 年版准确率自述:「In 75-80 per cent of the cases the examination correctly indicates the guilt or innocence of the accused; … 5 per cent or less is the margin of proven error.」——与第六章的 NAS 账对照着看,六十年口径几乎没动。
更新代际的普查结论同构:Shniderman(2012,Albany L.J. Sci. & Tech.[一手逐字])清点为 29 州 per se 禁、15 州约定可采、仅新墨西哥常态可采;State v. Sharpe(Alaska 2019,Justia 经 Wayback[一手逐字])转引初审法院对「all 50 states and the federal circuits」的调查:「30 jurisdictions still have a per se ban, 17 admit polygraph results based upon stipulation, and 12 leave the decision to the trial court’s discretion」。三个时点(2004/2012/2019)口径互差 2 州以内,结构稳定:约半数辖区全禁,约三分之一约定可采,常态可采的只有新墨西哥。
10.2 反向例外:新墨西哥
唯一的例外值得单独看,因为它是「反向红跳」的法律形态。Lee v. Martinez ¶4 holding(逐字):
“Instead, we hold that polygraph examination results are sufficiently reliable to be admitted under Rule 11-702, provided the expert is qualified and the examination was conducted in accordance with Rule 11-707.”
¶48 结论(逐字,本篇反向红跳在法律侧的对位句):
“…any doubt about the admissibility of scientific evidence should be resolved in favor of admission. … The remedy for the opponent of polygraph evidence is not exclusion; the remedy is cross-examination, presentation of rebuttal evidence, and argumentation.“
——不是排除,是交叉询问。50 州里唯一一个把「对抗制」而不是「排除制」当作测谎证据答案的辖区。Rule 11-707 附带严格条件(检查员资质、五年以上经验、全程录音录像)。它存在本身就是裁决的一部分:「全禁」与「全采」都不是美国法律的真实状态,真实状态是拼布加一块孤例外。
10.3 fMRI 出庭第一案:Semrau 2012
测谎的下一代技术第一次闯关,死在同一个门口。United States v. Semrau, 693 F.3d 510 (6th Cir. 2012)(第六巡回官方 PDF[一手逐字]):被告 Semrau 博士(被控医保欺诈)提交 Cephos 公司的 fMRI 测谎结果自证清白——「a matter of first impression in any jurisdiction」,全国首例。上诉法院一致维持排除,双重依据:
- Rule 702:「the technology had not been fully examined in ‘real world’ settings and the testing administered to Dr. Semrau was not consistent with tests done in research studies」;且「only Dr. Semrau knows whether he was lying … so there is no way to assess with complete certainty the accuracy of the two results finding he was ‘not deceptive’」——真值不可得的个案,正确率原则上不可评估。
- Rule 403(独立排除依据):检方事先不知情、宪法层面对「用测谎增强证人可信度」的警惕、结果无法对应具体指控陈述。
细节里的承重句(逐字):Cephos 创始人 Laken 在 Daubert 听证中先报各研究「accuracy rates between eighty-six percent and ninety-seven percent」,交叉询问时承认其 2009 年研究准确率「unexpected」跌至 71%,并承认存在「a huge false positive problem」——说真话者被误判为说谎「around sixty to seventy percent of the time」,某项研究中「nineteen out of twenty people that were telling the truth we would call liars」。第三个测试的补做(前两个一真一假)被法官形容为「begs the question whether a fourth scan would have revealed Dr. Semrau to be deceptive again」——重测到出想要的结果为止,这在测谎原生技术上做不到的事,在脑成像版本上第一天就出现了。
脚注 12 是法院留给未来的话(逐字):「The prospect of introducing fMRI lie detection results into criminal trials is undoubtedly intriguing and, perhaps, a little scary. … There may well come a time when the capabilities, reliability, and acceptance of fMRI lie detection—or even a technology not yet envisioned—advances to the point that a trial judge will conclude … Though we are not at that point today…」
后续:联邦司法中心《科学证据参考手册》第四版(2025,NCBI Bookshelf 全文[一手逐字])确认:fMRI 测谎证据至今未在任何法院被采纳,「we find no reported cases about such attempted uses since 2012」——2012 年后连尝试都没有了。
10.4 中国:一份批复,一张封闭清单
中国法域的规范地位由两份文件构成,均为一手:
① 最高人民检察院 1999 年批复(高检发研字〔1999〕12 号,1999-09-10,烟台市纪委监委官网法规库转载[一手逐字]),全文逐字:
“四川省人民检察院: 你院川检发研〔1999〕20号《关于CPS多道心理测试鉴定结论能否作为诉讼证据使用的请示》收悉。经研究,批复如下: CPS多道心理测试(俗称测谎)鉴定结论与刑事诉讼法规定的鉴定结论不同,不属于刑事诉讼法规定的证据种类。人民检察院办理案件,可以使用CPS多道心理测试鉴定结论帮助审查、判断证据,但不能将CPS多道心理测试鉴定结论作为证据使用。 此复”
「可以帮助审查、判断证据,但不能作为证据使用」——与 Lykken 的「侦查工具有用、法庭证据不行」立场结构同构,写成了一句公文。
② 刑诉法证据种类封闭清单(2018 版第五十条,与 2012 版第四十八条内容相同):八类证据(物证;书证;证人证言;被害人陈述;犯罪嫌疑人、被告人供述和辩解;鉴定意见;勘验、检查、辨认、侦查实验等笔录;视听资料、电子数据)中无心理测试结论。
任务书预设勘误(附录 A 第一项,此处登记结论):「2013 年最高法刑诉法解释将心理测试排除出证据种类」的通行表述不准确。本轮对三个同期规范文本做了全文级阴性检索:法释〔2012〕21 号《刑诉法解释》全文 548 条、《人民检察院刑事诉讼规则(试行)》(2012)全文 708 条、现行《人民检察院刑事诉讼规则》(2019)全文——「心理测试/测谎/多道/CPS」关键词全部零命中(方法备案见附录 A;[多源检索])。实际排除机制是 1999 批复的明文否定+证据种类封闭清单的缺位,学理文献一致确认「目前只有 1999 年批复这一法律文件」。
应用史节点(检察日报 2007-11-21《测谎仪:不可不信,不可全信》,新浪转载页[一手逐字]):1991 年国产第一台 PG-1 型多参量心理测试仪问世+首期测谎员培训班;1998 年检察机关开始应用;2007 年 1 月 1 日起心理测试纳入人民检察院鉴定机构与鉴定人登记管理。规模样例(同文逐字):大连市检察院至 2007 年 7 月在近 400 个案例中运用、678 人次被测试;苏州市沧浪区检察院 2004–2007 年参与 46 起案件、测试 110 人。
10.5 日本、以色列、英国
- 日本(Matsuda, Ogawa & Tsuneoka 2019, Front. Psychiatry,作者单位=警察厅科学警察研究所,[一手逐字]):「Japan is the only country where the polygraph with the concealed information test (CIT) is widely applied to criminal investigations.」「In Japan, the CIT is the only polygraph application used in criminal investigations. The CQT is not currently used at all. About 100 polygraph examiners deal with about 5,000 cases per year.」最高法院 1968 年曾有采纳测谎证据的判例,但实务上结果「rarely dealt with in court」,主要作侦查评估工具。日本选了美国学界论证更有基础的那条路(CIT),而美国自己没走——这个对照本身就是提问账的制度化裁决。
- 以色列:CIT/GKT 现场研究的产地(Elaad & Ben-Shakhar 1991 作者单位逐字=「Division of Criminal Identification, Israel National Police Headquarters」);NAS 2003 点名测谎广泛使用的国家「notably, Israel, Japan, and Canada」[一手逐字]。但现场效度账不乐观:两项以色列现场 GKT 研究假阴性率 42%/20%(假阳性 2%/5%)(Ben-Shakhar & Elaad 2003 引,一手逐字)——实验室 d=2.09 到了现场漏掉五分之二。
- 英国:不作法庭证据,但写入假释监管。Offender Management Act 2007 s.28 现行合并文本[一手逐字]:「The Secretary of State may include a polygraph condition in the licence」——适用范围经历次修订覆盖性犯罪、恐怖主义犯罪(2021 年扩入)与家暴犯罪(2021 年扩入)。用途=合规监管工具,不是定罪证据。
10.6 法域层小结
法域层裁决:美国是拼布(约半数 per se 禁/约三分之一约定可采/新墨西哥孤例常态可采),联邦系统名义裁量、实操排除;fMRI 测谎唯一一次出庭被 702+403 双杀且此后再无尝试;中国走「可辅不可证」的批复路线;日本全面押注 CIT;英国把它改造成假释条件。没有任何法域把测谎结论当作「机器读出真话」来用——所有制度都只在它外面包了一层自己的用途壳,壳的厚度各不相同,内核全是「辅助」二字。
十一、名号与消费账:谁叫它「测谎仪」
11.1 「lie detector」是报界造词,不是发明者造词
这个名字的谱系本身就是一次传播学审计。三个可核实的锚点:
- 报界造词。Invention & Technology(American Heritage 旗下,2004 年冬季号,Jack Kelly)逐字(原文 OCR 断词如此):「He published his findings in 1917, and when NEWS papers got wind of his work, they coined the term lie detector. Marston became the technology’s first great promoter; his enthusiasm set rolling the modern history of the lie detector, and his zeal eventually tarnished the device’s scientific credentials.」(原文[一手逐字])。Smithsonian Magazine(2024)对 Larson 一侧同向:「In a frenzy of sensationalist reporting, the press dubbed Larson’s polygraph a ‘lie detector,’ and the Examiner swooned: ‘All liars, regardless of cleverness, are doomed.’」(原文[一手逐字])。
- 可核的最早报刊用例在 1922 年。Lepore(Yale Law Journal 2015,全文[一手逐字])的档案脚注给出华盛顿各报 1922 年夏的标题:「Lie-Detector Verdict Today」(Wash. Post, 1922-07-20)——配的引文是 Marston 对 Frye 测试的预告:「Dr. Marston made a test of Frye’s blood pressure yesterday. Frye stoutly maintains that he is innocent of the crime.」;以及「Lie Detector Said to Clear Dudding in Killing of Uncle 12 Years Ago」(1922-08-02)。原始报纸页面受 Chronicling America 反爬拦截未能直取(附录 B 登记),经 Lepore 档案脚注转引。
- 词频曲线吻合。Google Ngram(en-2019 语料):「lie detector」1921 年起出现非零读数,1930–1932 年陡增(1.8e-8 → 3.9e-8),峰值在 1964 年(Ngram 交互图[一手逐字])——「词在 1920 年代初进入印刷品、1930 年代初爆发」。
谁「造」了这个词无定论(三说并存:报纸 1917 年后造/报界 1921 年为 Larson 机器造/Marston 自造说仅见于二手转述)。但词源学结论可以下:可核的用例全部来自报纸,不来自任何一位技术发明者。
11.2 三个发明者都拒绝这个词
更硬的证据在发明者自己的态度里——三个人,三种拒绝方式:
- Marston:「William Marston is sometimes credited as the inventor of the lie detector, though he always said that he originated a technique, not a device.」(InTech 2004,同上[一手逐字])——他坚持自己开创的是一种「技术」,不是一台「机器」。
- Larson:「Larson, who preferred the term emotion recorder…」(Weiss, Watson & Xuan, “Frye’s Backstory”, J Am Acad Psychiatry Law 2014,全文 PDF[一手逐字])——他给机器的命名是「情绪记录仪」。
- Keeler:1964 年 DePaul Law Review 的行业综述逐字:「The man who, in 1926, perfected the polygraph that is almost universally used today, Leonarde Keeler, objects to the use of the term ‘lie detector.’ Professor Keeler, often called the father of the ‘lie detector,’ compares the polygraph to the tools of the physician, stressing: ‘No matter what diseased condition is being sought, the instrument is still called a stethoscope, not a T.B. detector, a pneumonia detector…’」(PDF 经 Wayback[一手逐字])——被称为「测谎仪之父」的人,反对「测谎仪」这个词:医生的仪器叫听诊器,不叫「肺结核检测器」。
名号层裁决:「测谎仪」这个名字是媒体给的,造机器的三个人一个都不要它。 一百年后行业里每个人都用这个词——名字赢了,命名者输了。
11.3 Marston 在 Frye 案中的真实角色:人是真的,证没作成
通行叙事「Marston 是 Frye 案的专家证人」不准确。档案核验(Lepore 2015;Weiss/Watson/Xuan 2014,均同上)拼出的真实角色链:
- 1922 年 6 月 10 日,辩护律师(上过 Marston 夜校课程「Legal Psychology」)把教授带进 DC 监狱,Marston 对 Frye 施测(Lepore 逐字:「And then, on June 10, they brought their professor to the D.C. jail to meet the defendant. … Marston asked Frye if he would submit to the use of the lie detector; Frye agreed.」)。Marston 1938 年自叙:「I gave him a deception test in the District jail. No one could have been more surprised than myself to find that Frye’s final story of innocence was entirely truthful!」
- 庭审中辩方 proffer 他为专家证人,主审法官 Walter McCoy 次日裁定不准作证,当庭演示的请求也被拒(Frye 判决原文:「The offer was objected to by counsel for the government, and the court sustained the objection. Counsel for defendant then offered to have the proffered witness conduct a test in the presence of the jury. This also was denied.」)。
- Frye 二级谋杀罪成立——不是「因测谎获释」。
- 两级法院判决均不点名 Marston(Lepore:「Marston’s name is not mentioned in the opinions of either the trial or the appellate court.」);上诉判决只说「an expert witness」。
- 当时报纸的标题(Weiss 等 2014 引):「Court Rules Out Lie-Finding Device」「Invention Met Its Death on First Trial」「Quick Death to ‘Sphygmomanometer’」——Weiss 等的观察逐字:「It appears that the fate of the lie detector sold more newspapers than stories of the underlying crime. Marston and his machine had become celebrities.」
即:测了、被提名、被拒、罪成、判决不点名——「人是真的,证没作成」。而 Marston 的自我营销账同时是真的:给 Lindbergh 绑架案嫌疑人提议施测、带着仪器拍吉列剃须刀广告(1938 年 Life 杂志)、在电台讲「如何服从冲动」(Bunn 1997,作者自存 PDF[一手逐字]);1923 年 3 月——Frye 判决前十个月——他自己因邮件欺诈被捕,报题「Marston, Lie Meter Inventor, Arrested」(Lepore 2015 脚注转引[需亲核])。测谎的第一个推销员,先成了新闻,再成了判例,最后成了广告。
11.4 电视与中文谱系:仪器隐身的两个样本
电视侧:Fox 真人秀 The Moment of Truth(2008-01-23 首播)——选手在镜头前回答 21 个问题赢 50 万美元,判定依据是开播前录制的一次测谎(USA Today 经 ABC News 转载[一手逐字])。节目已售往 26 国;哥伦比亚原版停播原因两说互斥(USA Today:女选手自白雇凶杀夫;Plain Dealer:参赛者自杀,且把「Colombia」拼成「Columbia」)——两个权威媒体给出互相矛盾的死因,本身就是媒体层讹传的活样本(附录 B 登记,两说并存不裁决)。
对这篇报告最有价值的是 Plain Dealer 的机制观察(原文[一手逐字]):「the real star of the show is the lie detector」——而「On ‘Moment of Truth,’ the ‘lie detector’ is never shown. A disembodied female voice renders the verdict.」——仪器隐身,只剩一个没有形体的女声宣判「True/False」。这就是名号跳的终端形态:连机器都不必出现,名号自己就能宣判。
中文侧:「测谎仪」三个字在官方话语里的标准用法,检察日报 2007 年那篇的标题就是——《测谎仪:不可不信,不可全信》;中新网 2009 年北京检方报道:「心理情景测试仪被俗称为测谎仪……其测试结果尚不能作为定案证据」,案例句「面对测试结果,他最终向检察官如实供述了全部受贿过程」(中新网[一手逐字])——注意这个因果句式:起作用的是「面对测试结果」,即威慑-诱供效用,与 NAS「distinct from actual validity or accuracy」的切割完全同构。人民网 2014 年郑州市检察院测谎师自述:「只是一种心理测试,有时会与测谎对象有智慧、阅历上的较量」、被测者「骗得了别人,骗不了自己」(人民网[一手逐字])——从业者口语里的「测谎」,承认是「较量」,不是「读数」。
11.5 市场账:三组对峙数字
消费层的最后一笔是钱与宣称。三组对峙全部一手:
① APA vs NAS。 行业协会官网自述:1966 年成立,「2700+ members」,「the largest polygraph association」;准确率承诺:「Through strict adherence to training and education standards, APA examiners are able to attain accuracy rates exceeding 90 percent.」(APA About 页[一手逐字])。对峙面是 NAS 2003 的裁决:「reliance on polygraph testing to perform in practical applications at a level at or above A = 0.90 is not warranted on the basis of either scientific theory or empirical data」(第 5 章[一手逐字])+「insufficient to justify reliance」(第七章)。同一片技术,招牌写 90%+,委员会写 0.86 且很可能高估。
② CVSA(语音压力分析):宣称 96.4%,实测「不比抛硬币好」。 厂商 NITV 官网逐字宣称:「in 96.4% of interviews conducted, where the CVSA® indicated stress, suspects made self-incriminating confessions」(NITV 官网[一手逐字])。对照组三份一手文件:NIJ 资助的现场研究(Damphousse 等,2007,OJP 官方 PDF[一手逐字])在县监狱用尿检做真值:「Both VSA programs show poor validity … The programs were not able to detect deception at a rate any better than chance」,且专家与新手解读一致性仅 0.11–0.52;弗吉尼亚州职业监管委员会 2003 年裁定「The studies have produced no evidence that the use of the CVSA provides accuracy rates better than chance.」并记录 DoDPI 1996 盲评总准确率 49.8%、「not significantly different from chance」(州文件 PDF,APA 官网托管[一手逐字])——行业协会自己托管了证伪文件,这个细节值得留着。
③ fMRI 双公司兴衰:从「年底上市」到全网 404。 2006 年 1 月,WIRED 当期报道(原文[一手逐字]):「By the end of 2006, two companies, No Lie MRI and Cephos, will bring fMRI’s ability to detect deception to market.」Cephos 的 Laken 说「FMRI lie detection is where DNA diagnostics were 10 or 15 years ago」;No Lie MRI 的 Huizenga 说要建「VeraCenters」网络,同时承认「this is a company – we’re here to make money」;CBS/AP 同月(原文[一手逐字])记下 O.J. 案辩护律师 Robert Shapiro(Cephos 顾问且有经济利益)的话:「I’d use it tomorrow in virtually every criminal and civil case on my desk.」No Lie MRI 2006 年官网首页(Wayback 2006 快照[一手逐字])把目标市场列成联邦清单:国防部、国土安全部、司法部、CIA。结局:2009 年初 San Diego 抚养权案采纳动议撤回(No Lie MRI;FJC 手册第三版神经科学章逐字记录);2010 年 Wilson v. Corestaff 按纽约州 Frye 排除(Cephos);2012 年 Semrau 第六巡回维持排除(10.3);Wayback CDX 时间线显示 noliemri.com 自 2020 年 11 月起连续 404、2023 年后恢复的快照为注册商停放页。从「给每个案子都用上」到域名停放,六年。 措辞节制:无公司终止运营的正式声明,可坐实的表述是「网站下线、法庭连败」。
11.6 消费层小结
名号与消费层裁决:词是报纸造的,机器的三位造者都拒收;第一推销员把自己卖成了新闻再卖成了判例;电视时代连仪器都可以隐身,只剩宣判的嗓音;市场宣称(90%+/96.4%)与官方实测(0.86 且很可能高估/不比随机好)之间的缝,从来不是数据之争,是生意与证据各自的分工。
十二、反措施账:约半数能破,教这件事的人入狱
反措施(countermeasures)是测谎的阿喀琉斯之踵:如果受测者能人为操控对照题上的反应,CQT 的整个比较逻辑就塌了。这一章的账分两侧:实验室证明能不能破,官方回答能不能查,以及——教别人破的人后来怎么样了。
12.1 攻击侧:约 50%,且「仪器和观察都查不出」
核心实证是 Honts, Raskin & Kircher 1994(J. Applied Psychology 79(2):252–9,PubMed 摘要逐字[一手逐字]):120 名社区受试者、现场技术施测;身体反措施(咬舌、脚趾压地)与心理反措施(倒数 7)等效:
“The mental and physical countermeasures were equally effective: Each enabled approximately 50% of the Ss to defeat the polygraph test. … Moreover, the countermeasures were difficult to detect either instrumentally or through observation.”
——约半数受训者击败测试,且反措施本身「仪器上和行为观察上都难以检测」。
对 CIT 的效应是混杂的,不是免疫也不是免疫破解(三篇摘要逐字):
- Honts 等 1996(Psychophysiology 33(1),PubMed[一手逐字]):「the CKT has no special immunity to the effects of countermeasures」——CIT 同样被降级。
- Ben-Shakhar & Dolev 1996(J. Applied Psychology 81(3),PubMed[一手逐字]):反措施条件下「a significant reduction in electrodermal detection efficiency」。
- Elaad & Ben-Shakhar 1991(Int. J. Psychophysiology 11(2),PubMed[一手逐字]):「the item-specific countermeasures tended to increase psychophysiological detection, whereas the continuous dissociations tended to decrease detection efficiencies」——项目特异反措施反而提高了检测率。反措施不是「一用就灵」的咒语,用错方式会把自己标得更亮。
另有一条不可核实级线索登记在此:AntiPolygraph.org 2018 年博客转述一项「secret 1995 study」,称「80 percent of test subjects who were taught polygraph countermeasures succeeded in beating the U.S. government’s primary counterintelligence polygraph screening technique. The training took no more than an hour.」(倡导站点博文[需亲核])——研究原件据称涉密、从未公开,只能标注为「倡导站点转述的未公开研究」,不承重。
12.2 防守侧:NCCA 的立场(经泄露件)
检测方立场的一手来源是一份经 AntiPolygraph.org 公开的 NCCA 内部白皮书(「Timeline Detailing the Countermeasures Classification Issue」,2012 年,泄露件 PDF[一手逐字];泄露件性质先标注:内容代表 NCCA 自我立场陈述,真实性未经 NCCA 确认):
“NCCA personnel are officially and unofficially considered the Subject Matter Experts (SMEs) for defining and identifying polygraph countermeasures in the federal government.”
“the use of polygraph results has become the most viable tool available for use by federal LE agencies to disallow applicants from being hired into sensitive positions. NSA and CIA have successfully use the polygraph in this manner for years in a classified environment. With the sophistication of CM procedures, the identification of persons employing CMs is also being effectively used for denying persons access to sensitive LE positions.”
第二句值得读两遍:官方立场不是「我们能保证测出谎言」,而是「识别使用反措施的人,本身就是拒绝录用的工具」——测不出你是否撒谎,但可以因为你试图影响测试而拒绝你。防御的真实形态是把「对抗测试」本身定罪化。(文件另有一处数量级错误:「since 2000 the federal government has conducted over 1,000,000,000 examinations」——10 亿次明显有误,与全部官方规模账差三个数量级,引用时标注。)
12.3 NAS 的居中裁决:unknown
委员会对攻与防的最终表述(执行摘要,逐字):
“Certain countermeasures apparently can, under some laboratory conditions, enable a deceptive individual to appear nondeceptive and avoid detection by an examiner. It is unknown whether a deceptive individual can produce responses that mimic the physiological responses of a nondeceptive individual well enough to fool an examiner trained to look for behavioral and physiological signatures of countermeasures.”
“CONCLUSION: Basic science and polygraph research give reason for concern that polygraph test accuracy may be degraded by countermeasures, particularly when used by major security threats… If these measures are effective, they could seriously undermine any value of polygraph security screening.”
实验室里能破(已知);能否骗过受过反措施识别训练的检查员(未知);而对筛查场景——恰恰是最需要担心反措施的场景——「如果有效,可能严重破坏测谎安全筛查的任何价值」。委员会把最重的担忧留给了国家行为体:major security threats 恰是有资源训练反措施的对象。
12.4 教反措施的人:Operation Lie Busters、Williams、Dixon
法律对「教反措施」的处置,是这篇报告里最新也最锋利的一页。
行动与定性:McClatchy 2013-08-16 调查报道(Wayback 快照[一手逐字]):联邦特工对「声称能教求职者通过测谎」的培训师启动刑事调查,属奥巴马政府「unprecedented crackdown on security violators and leakers」的一部分;CBP 的专项代号「Operation Lie Busters」,10 名申请人因试图使用反措施被拒录。
Doug Williams 案(一手两份 DOJ/FBI 新闻稿):Williams——前俄克拉何马城警方测谎员,公开教授反措施三十年,上过 60 Minutes 等全国性节目,且(McClatchy 逐字)「testified in congressional hearings that led to the 1988 banning of polygraph testing by most private employers」——他正是当年作证促成 EPPA 的人。2014-11-14 被起诉(W.D. Okla.,5 项罪名);2015-05-13 认罪(2 项邮件欺诈+3 项妨害证人,FBI 新闻稿[一手逐字]);2015-09-22 判刑两年(DOJ 新闻稿,经 Wayback[一手逐字])。案情核心:他训练一名假扮联邦执法官员的卧底「to lie and conceal involvement in criminal activity」,并教学员「deny receiving his polygraph training」。调查方是 CBP 内务办公室+FBI。
Chad Dixon 案(DOJ EDVA 新闻稿,经 Wayback[一手逐字]):受 Williams 著作启发的印第安纳培训师;2012-12-17 认罪(电汇欺诈+妨害机构程序);2013-09-06 判 8 个月监禁+3 年监外监管+没收 17,091.07 美元;收费每次 1,000–2,000 美元;客户逐一点名:持 Top Secret 许可的联邦承包商、CBP 职位申请人、九名被判刑的性犯罪者;第二名卧底自称与未成年人发生性关系、其兄为 Los Zetas 成员,Dixon 仍施训并嘱其隐瞒。
12.5 反措施层小结与讽刺格
反措施层裁决:攻击侧实证(约半数可破 CQT、难检测;对 CIT 效应混杂)与防守侧立场(识别反措施本身作为拒绝理由)并存,委员会居中落「unknown」。而法律给出的答案不是技术改进,是刑罚:三十年前作证促成「禁止私人测谎」的人,三十年后被政府测谎体系以「教人通过测谎」送进监狱。当年他用国会听证削弱了私人部门的测谎,如今联邦用联邦监狱保护了政府自己的测谎——同一个人,两部法律,方向相反。这个讽刺格不在任何人的计划里,它是制度演化自己写出来的。
十三、反向红跳:四句流行的「反转」也逐句裁决
对称红队的规矩:拆完正向跳变,必须回头拆反向跳变——那些「既然测谎不灵,那就可以说……」的句子。四句,逐句裁决。
13.1 「测谎和抛硬币一样」→ 不立
这是传播链上最常见的反转句,它不立。裁决依据全部是官方文件逐字:
- 「specific-incident polygraph tests can discriminate lying from truth telling at rates well above chance, though well below perfection」(NAS 2003 执行摘要)。
- 「the data … clearly fall above the diagonal line, which represents chance accuracy. Thus, we conclude that features of polygraph charts and the judgments made from them are correlated with deception」(第 5 章)。
- 中位 A=0.86(52 个实验室数据集);CIT 元分析 d=1.55、模拟犯罪子集 d=2.09(B&E 2003)。
- OTA 1983 同向:「detects deception at a rate better than chance, but with error rates that could be considered significant」。
限定同样写死:「好于随机」是个弱命题——随机水平是 50%,司法定论与国家安全筛查的门槛远高于此;且批评阵营最强一手(Iacono 2002 参院证词)的「little better than chance」是限定在无辜者被判有罪方向(特异性)与现场 CQT 上的:「innocent people fare little better than chance on these tests, with 40% or more failing on average」(参议院听证记录[一手逐字])。正确的表述是:总体判别显著好于随机,但「好于随机」撑不起它被要求的任何工作;在低基数筛查场景,好于随机仍然不够用。
13.2 「反措施傻瓜可破」→ 部分为真,被夸大
这一句有真实的核:Honts 1994 的约 50% 击败率、「difficult to detect either instrumentally or through observation」、NAS 承认实验室条件下反措施可降级准确率——都是一手逐字(第十二章)。但「傻瓜可破」超出证据四个身位:
- 成功率是约一半,不是必然;且需要训练,不是看半小时网页。
- 对 CIT 的效应混杂:Honts 1996 说无免疫,Elaad & Ben-Shakhar 1991 说项目特异反措施反而提高检测——用错方法会自我暴露。
- 官方层面是 unknown:「能否骗过受过反措施识别训练的检查员」NAS 的答案是不知道,不是「能」。
- 制度已把「使用反措施」本身定罪化:NCCA 白皮书立场+Operation Lie Busters+Williams/Dixon 两案——教与用都是刑事风险。另外那条「1995 年研究 80% 击败率」的说法只存在于倡导站点转述,原件据称涉密未公开,不可核实,不得用来承重。
13.3 「美国法律全禁」→ 不立
EPPA 1988 禁的是私人雇主;§2006(a) 明文豁免联邦、州、地方政府,(b) 点名豁免 NSA/DIA/NGA/CIA;CBP 依 2010 年法律法定强制全员测谎(第七章)。刑事法庭上:约半数辖区 per se 排除、约三分之一约定可采、新墨西哥常态可采(第十章)。正确表述:私人入职筛查被禁,政府与国家安全领域不仅合法而且制度化强制;刑事可采性是拼布加一块孤例外。
13.4 「脑成像已解决测谎」→ 不立
fMRI 测谎 2006 年宣布「年底上市」,2012 年唯一一次出庭被 702+403 双重排除,此后连尝试都没有(FJC 2025:「no reported cases about such attempted uses since 2012」)。科学侧的裁决同样清楚:Monteleone 等 2009(Soc. Neurosci.,PubMed[一手逐字])——「no region could be used to correctly detect deception across all individuals」,最好的脑区(内侧前额叶)个体水平只到 71%,标题本身就是「better than chance, but well below perfection」;领域创始人之一 Langleben 与 Moriarty 2013 对 Semrau 排除的回应是「a point with which we concur」(PMC 全文[一手逐字])。「已解决」无任何一手支撑。
13.5 反向层小结
四句反转,三句不立,一句(反措施)有核但被夸大。反向红跳的共同形状与正向跳变同构:都是把一个有严格限定词的结论,砍掉限定词再流通。「well above chance, though well below perfection」砍掉后半句是行业广告,砍掉前半句是虚无口号——委员会写的是一整句,传播链从来只取半句。
十四、对称金句与三向红线
14.1 对称金句
机器记录的是唤醒,宣判的是人。
一百年来被测试的从来不是谎言——是一个制度愿意为多低的基数、错标多少清白者。那台机器给美国法律留下了统治科学证据七十年的判例,而它自己,至今不被允许出庭作证。名字是报纸给的,造它的三个人一个都不要;抓间谍是它的正当性叙事,而三宗最大的间谍案没有一个由它检出;教别人通过它的人进了监狱,而它的考官每年对十一万人宣判「欺骗」或「清白」。这台机器最真实的产物,从来不是读数,是制度围绕读数做出的那些决定。
14.2 三向红线(本篇结论不允许的三向读法)
- 不许读成「测谎有效,放心推广」:信号无特异(4.2)、理论缺席(5.4)、个案效度系统性高估(第六章)、筛查场景被委员会明判不可用(第七章)、三宗间谍大案无一检出(第八章)。
- 不许读成「测谎是骗术,一律废除」:个案判别显著好于随机(守真锚二)、CIT 有真实科学基础与日本全国实务(守真锚三)、威慑-诱供效用被官方承认(守真锚五)、批评阵营领袖 Lykken 本人承认个案用途(守真锚六)、新墨西哥以对抗制而非排除制接纳它(10.2)。
- 不许读成「技术会自己解决」:fMRI 版本唯一一次出庭被双杀且此后再无尝试(10.3);个体水平最高约七成、无现场错误率(13.4);「intriguing and, perhaps, a little scary」的那扇门至今仍关着。
14.3 自指:本篇自己的承重方式
本篇的结论结构是「一个记录情绪唤醒的仪器在受控提问下与欺骗存在统计关联(真)——被读成一台能从人体内读出真话的机器(假)」。这个结构本身的承重点在三处:NAS 2003 与 OTA 1983 的官方文本(逐字引用最多)、间谍案档案(SSCI/OIG/FBI 官方页)、判例原文(Frye/Daubert/Scheffer/Lee/Semrau)。凡这三处之外的承重(行业宣称、媒体叙事、泄露件、倡导站点转述),本篇一律降格标注或移入附录 B。本篇最强的限定词与最强的事实来自同一份文件——这是「对称」二字在本篇的具体含义:拆机器的话术,用的是委员会自己的算术;挡虚无的口号,用的还是委员会自己的算术。
附录 A:任务书预设勘误明细(6 处)
- 「2013 年最高法刑诉法解释将心理测试排除出证据种类」——证伪。 三个同期规范文本全文级阴性检索:法释〔2012〕21 号(548 条,新旧条文对照镜像 GBK 全文 182,648 字符,另最高法公报官网版取回文首与全部目录)、《人民检察院刑事诉讼规则(试行)》(2012,708 条镜像全文)、《人民检察院刑事诉讼规则》(2019,官方 PDF 97 页)——「心理测试/测谎/多道/CPS」关键词全部零命中[多源检索]。实际排除机制=高检发研字〔1999〕12 号批复明文否定+刑诉法证据种类封闭清单缺位;学理文献(如《中国法学》2024 年第 1 期郑飞文)一致确认「目前只有 1999 年批复这一法律文件」。
- Lee v. Martinez 案号:任务书预设「2004-NMSC-051」,正确为 2004-NMSC-027(CourtListener cluster 2623542,三并列引证一致)。
- NAS 自算表的 A 值:池表写「AUC 约 0.85」;报告自算表统一用 A=0.90 且明言其为高于实测的乐观上界,0.85 级数字对应实测中位 0.86(IQR 0.81–0.91)。两者不可混用。
- 「Marston 是 Frye 案专家证人」:不准确。准确表述=施测(1922-06-10 监狱)→辩方提名→McCoy 法官拒其作证→Frye 二级谋杀成立→两级判决不点名(11.3)。
- 「一页纸」表述:按可核计量修正为「全文 669 词(Lepore 2015)/Federal Reporter 占 1013–1014 页,不足两页」。
- Iacono & Lykken 1997 精确百分比:原文付费墙未破(附录 B 第一条),本篇只承接到「约三分之一认为 CQT 科学上成立/约四分之一认为应可采」级,经同一第一作者 2001 年复述与 FJC 手册第三版双源交叉(6.4)。
附录 B:取不回清单(缺席登记,缺席不等于不存在)
- Iacono & Lykken 1997(J. Applied Psychology 82(3))原文:APA 付费墙;PsycNet 的 Wayback 快照为 JS 壳;Semantic Scholar 摘要被出版商隐藏。SPR 调查精确百分比未获一手逐字。
- Lykken 1959(J. Applied Psychology,GKT 原始论文)原文:付费墙,本轮未破。
- Honts, Thurber & Handler 2021 元分析全文:付费墙;效应量 0.69 仅摘要级(OpenAlex 元数据[需亲核])。
- Saxe 1985 等早期 OTA 引用研究原始摘要:未逐一取回。
- DoD FY2011 年报 PDF 原件(43,434 次/41,057 筛查次):仅 APA Magazine 与 ClearanceJobs 二手转引;两轮检索未找到原件,建议后续走 NCCA/DCSA FOIA。
- Giannelli 1980(80 Colum. L. Rev. 1197)原文:bepress PDF 对直取返空、Wayback 无快照;引用史数字依赖 Faigman 1994 fn.25 转述(其注明源自 Giannelli 1983 与 Starrs 1982)。
- Semrau W.D. Tenn. 2010 R&R 原件(2010 WL 6845092):RECAP 文书不可用、Westlaw 付费墙;关键推理已由第六巡回判决逐字转引+FJC 2025 指南兜底。
- Chronicling America(loc.gov 报纸档案):Cloudflare 拦截两轮,1922 年报纸原始版面未直取;最早报刊用例经 Lepore 2015 档案脚注转引。
- 正义网/检察日报老站原文:404 模板与 JS 壳;《南方日报》2009 报道 PDF 字体不可抽取。中文样例依赖中新网、人民网与烟台政府网转载件。
- NCCA 1995 年「80% 击败」研究原件:据称涉密未公开,仅倡导站点转述,不承重。
- DOE 1999 规则覆盖人群总数一手估计(二手常称约 13,000–20,000):本轮未找到一手出处,未采用。
- EPPA 立法史众议院报告原件(1988 年前私人部门年测谎量):未取。
- 哥伦比亚版 The Moment of Truth 停播原因一手声明:两说互斥(雇凶杀夫自白/参赛者自杀),未裁决。
- NIJ 2012 年 CVSA 现场评估报告(Chapman 参与)原件:未取回;「同一学者两边出现」的精确表述待核对。
- Polygraph Rules 2009(英国 SI 2009/619)原件:未取;OMA s.28 合并文本已足够支撑主命题。
附录 C:来源清单(全部为本轮实际取回并落盘件;Wayback 链接为 id_ 原件快照)
核心官方报告
- National Research Council (2003), The Polygraph and Lie Detection, National Academies Press, DOI 10.17226/10420:执行摘要/第 1 章/第 2 章/第 3 章/第 4 章/第 5 章/第 7 章/结论章
- OTA (1983), Scientific Validity of Polygraph Testing, OTA-TM-H-15:Princeton 镜像全文 PDF;FAS 章节镜像
- 参议院司法委员会听证会(2002,107th Cong.):govinfo 全文
- 联邦司法中心《科学证据参考手册》第三版(2011):FJC 官方 PDF;第四版(2025)神经科学指南:NCBI Bookshelf
判例
- Frye v. United States, 293 F. 1013 (D.C. Cir. 1923):CourtListener 全文经 Wayback
- Daubert v. Merrell Dow, 509 U.S. 579 (1993):Cornell LII
- United States v. Scheffer, 523 U.S. 303 (1998):多数意见/Stevens 异议
- Lee v. Martinez, 2004-NMSC-027 (2004):FindLaw 全文
- State v. Valdez, 91 Ariz. 274 (1962):CourtListener 经 Wayback
- State v. Sharpe (Alaska 2019):Justia 经 Wayback
- United States v. Semrau, 693 F.3d 510 (6th Cir. 2012):第六巡回官方 PDF
- Shniderman (2012), 22 Alb. L.J. Sci. & Tech. 433:全文 PDF
- Faigman, Porter & Saks (1994), 15 Cardozo L. Rev. 1799:UC Hastings 库经 Wayback
- Lepore (2015), Yale Law Journal「On Evidence」:全文
- Weiss, Watson & Xuan (2014), J Am Acad Psychiatry Law 42:226:全文 PDF
筛查制度与间谍档案
- DoD FY1999 致国会年报(Polygraph 29(3) 重刊):APA 期刊 PDF
- DCSA FY2026/FY2024 预算论证(NCCA 节):FY2026/FY2024
- DOJ OIG (2006), Use of Polygraph Examinations in the Department of Justice:Wayback 快照
- DOJ OIG (2003), Hanssen 案审查解密执行摘要:Wayback 快照;OIG (2007) 复查:theblackvault 镜像
- SSCI (1994), Ames 案评估:FAS 全文镜像
- FBI 官方案例页 Ana Montes:fbi.gov
- MI5 官方历史页 Klaus Fuchs:Wayback 快照
- EPPA 1988(29 U.S.C. ch.22):govinfo 法典
- Anti-Border Corruption Act 2010(P.L. 111-376):govinfo
- DOE 1999 终规则(64 FR 70962):联邦公报全文;DOE 2006 终规则(71 FR 57386):联邦公报全文
- GAO-19-419T(CBP 招聘证词,2019):GAO
- BJS《Local Police Departments, 2007》(NCJ 231174):PDF 镜像
学术研究与元分析
- Ben-Shakhar & Elaad (2003), J. Applied Psychology 88(1):作者自存全文;另章节自存稿
- APA 元分析(2011, Polygraph 40(4)):官方 PDF
- Nelson (2015), 测谎科学基础综述(行业侧):APA 官方 PDF
- Honts 等 (1994),J. Applied Psychology 79(2):PubMed;(1996) Psychophysiology 33(1):PubMed
- Ben-Shakhar & Dolev (1996),J. Applied Psychology 81(3):PubMed
- Elaad & Ben-Shakhar (1991),Int. J. Psychophysiology 11(2):PubMed
- Monteleone 等 (2009),Soc. Neurosci.:PubMed
- Langleben & Moriarty (2013):PMC 全文
- Honts, Thurber & Handler (2021):DOI 元数据(全文付费墙,附录 B)
- Iacono (2001)(含 I&L 1997 复述):镜像全文
- Matsuda, Ogawa & Tsuneoka (2019),Front. Psychiatry 10:24:全文
中国法域与中文话语
- 最高检 1999 批复(高检发研字〔1999〕12 号):烟台市纪委监委官网法规库
- 检察日报(2007-11-21)《测谎仪:不可不信,不可全信》:新浪转载
- 中新网(2009-11-27)北京检方测谎报道:中新网
- 人民网(2014-04-21)郑州测谎师报道:人民网
名号、流行文化与市场
- Kelly (2004), Invention & Technology:原文
- Smithsonian Magazine (2024) Larson 史:原文
- Lampert (1964), 13 DePaul L. Rev. 287:PDF 经 Wayback
- Bunn (1997), History of the Human Sciences 10(1):作者自存 PDF
- Google Ngram「lie detector」:交互图
- Larson 档案指南(OAC):典藏页
- USA Today(2008-01-23)The Moment of Truth:ABC News 转载;Plain Dealer(2008-01-28):原文
- APA 官网:About 页
- NITV/CVSA 官网:对照页
- NIJ (2007) VSA 监狱现场研究:OJP 官方 PDF;Virginia BPOR (2003):APA 托管 PDF
- WIRED (2006-01) fMRI 测谎:原文;CBS/AP (2006-01-30):原文;No Lie MRI 2006 首页:Wayback 快照
反措施与执法处置
- Honts 1994/1996、Ben-Shakhar & Dolev 1996、Elaad & Ben-Shakhar 1991(见上「学术研究」)
- NAS 2003 反措施结论:执行摘要(见上)
- NCCA 反措施白皮书(泄露件,2012):AntiPolygraph.org 托管 PDF;「1995 研究」转述:倡导站点博文(均按泄露件/转述级标注)
- McClatchy(2013-08-16)Operation Lie Busters:Wayback 快照
- Williams 案:FBI 认罪稿(2015-05-13)fbi.gov;DOJ 量刑稿(2015-09-22)Wayback 快照
- Dixon 案:DOJ EDVA 量刑稿(2013-09-06)Wayback 快照
其他
- Lykken 立场与「盾牌」论:FAS 2006 年纪念文(全文)
- 英国 Offender Management Act 2007 s.28:legislation.gov.uk 现行合并文本
本篇为 Chat Research 机制裁决系列第 151 篇(对称双向第 146 篇,section H,全库第 209 篇)。调研与写作:Kimi。