Calibrated Human-AI Trust: Why 'More AI' Isn't Better, and What a Deference-Calibrated AI Looks Like
校准的人机信任:为什么“更多 AI”不一定更好,以及“退让被校准的 AI”长什么样
A 2024 *Nature Human Behaviour* meta-analysis of 106 experiments found that putting a human and an AI together usually performs *worse* than the better of the two alone -- especially on decisions. The lesson for KKMatch: durable trust comes from *calibrated deference*, not more automation. We turn the evidence into a falsifiable hypothesis (KH-020) and a first-party experiment.
2024 年《自然·人类行为》一项涵盖 106 个实验的元分析发现,把人与 AI 放在一起,通常比两者中更强的那一个*更差*——尤其在决策任务上。对 KKMatch 的启示:持久的信任来自*校准的退让*,而非更多自动化。我们把证据转化为可被证伪的假设(KH-020)与一项第一方实验。
KKMatch Human Intelligence Research TeamKKMatch 人类智能研究团队· Research Lead: KK Research· Published: 2026-09-26· Reviewed by: KK Research· 10 min read
Executive Summary
执行摘要
In 2024, Vaccaro, Almaatouq and Malone published the first large-scale meta-analysis of human-AI collaboration in Nature Human Behaviour (106 experiments, 370 effect sizes). The headline result upends the 'AI everywhere' assumption: on average, human-AI combinations performed significantly worse than the best of humans or AI alone (Hedges' g = -0.23). The loss was concentrated in decision-making tasks (g = -0.27, p = 0.002); the gain appeared in content-creation tasks (g = +0.19). A 2023 field experiment at Boston Consulting Group (Dell'Acqua et al., 758 consultants) found the same shape: AI helped on tasks inside its capability frontier (+12% tasks, +25% speed, +40% quality) but hurt on a task just beyond it (19 percentage points fewer correct answers with AI). For KKMatch -- a relationship-intelligence platform built on an 11-dimension Human Model -- the takeaway is not 'use less AI.' It is that durable human-AI (and human-human) trust depends on calibrated deference: knowing, per task, whether to defer to the AI, to the human, or to hand control back. We formalize this as a new KK hypothesis (KH-020) and a first-party experiment, and we separate what the evidence proves from what remains KK's own, unvalidated claim.
2024 年,Vaccaro、Almaatouq 与 Malone 在《自然·人类行为》发表了首项大规模人机协作元分析(106 个实验、370 个效应量)。其标题式结论颠覆了“AI 无处不在”的假设:平均而言,人机组合的表现显著差于人或 AI 单独中的更强者(Hedges' g = -0.23)。损失集中在决策任务(g = -0.27,p = 0.002),收益出现在内容创作任务(g = +0.19)。2023 年波士顿咨询公司的一项现场实验(Dell'Acqua 等,758 名顾问)呈现同样形态:AI 在其能力边界之内有帮助(+12% 任务量、+25% 速度、+40% 质量),但在边界之外的一个任务上有害(用 AI 后正确答案少 19 个百分点)。对以 11 维 Human Model 为基石的关系智能平台KKMatch 而言,启示不是“少用 AI”,而是:持久的人机(以及人人)信任取决于校准的退让——按任务判断该让渡给 AI、让人、还是交还控制权。我们将其形式化为新的 KK 假设(KH-020)与一项第一方实验,并区分证据已证与仍属 KK 自身、未验证的主张。
Human-AI combinations vs the better of the two alone (Vaccaro et al., 2024)
人机组合 vs 两者中更强者(Vaccaro 等,2024)
Hedges' g from 106 experiments / 370 effect sizes. Green = the pair beat its baseline; red = the pair underperformed the best of human or AI alone ('no synergy'). Decision tasks lost most; creation tasks gained (but not significantly).来自 106 个实验 / 370 个效应量的 Hedges' g。绿色=组合优于其基线;红色=组合弱于人或 AI 单独中的更强者(“无协同”)。决策任务损失最大,创作任务受益(但不显著)。
Source: Vaccaro, Almaatouq & Malone (2024), Nature Human Behaviour 8(12):2293-2303, DOI 10.1038/s41562-024-02024-1.
The meta-analysis lands exactly where KK's Human Model already points. KH-002 (trust formation can be decoupled from absolute answer quality), KH-005 (personalization is changing interaction strategy, not knowing more), and KH-006 (long-term quality depends on Interaction Adaptation) all predict that 'more AI' is not the unit of value -- how the system hands control back and forth is. The decision-task loss is, we argue, a failure of calibrated deference: when the AI is more capable, people still override or fail to catch its errors; when the human is more capable, the AI drags them down. For KKMatch, this reframes both product and science. On the product side, an AI companion that models each user's calibration discernment relative to their Personal Baseline (KH-001) -- when this user tends to over-trust or under-trust -- and that signals its own uncertainty on tasks where the human holds private context, is a concrete instantiation of KH-005/KH-014. On the science side, the same mechanism operates between two humans: durable rapport in a relationship depends less on 'how much each partner knows' and more on whether each can read when to defer, challenge, or step back -- which is precisely the Dyadic Model of KH-004 and the repair logic of KH-013. The 11-dimension Human Model gives KK a place to measure calibration discernment as a trait+state, not just assert it.
元分析的结论正好落在我 KK 的 Human Model 早已指向之处。KH-002(信任形成可与绝对回答质量脱钩)、KH-005(个性化是改变交互策略而非知道更多)、KH-006(长期质量取决于交互适应)都预测:“更多 AI”并非价值单位——系统如何来回交还控制权才是。我们认为,决策任务的损失是一种校准退让的失败:当 AI 更强时,人仍会推翻它或漏掉它的错误;当人更强时,AI 反而把人拖下水。对 KKMatch 而言,这重新框定了产品与科学两面。产品侧,一个 AI 伴侣若相对于用户个人基线(KH-001)建模其“校准判断力”——此人倾向于过度信任还是欠信任——并在用户握有私有情境的任务上暴露自身不确定性,正是 KH-005/KH-014 的具体落地。科学侧,同样机制也发生在两人之间:关系中的持久融洽,较少取决于“每位伴侣知道多少”,更多取决于彼此能否读懂何时退让、质疑或退后——这正是 KH-004 的二元模型与 KH-013 的修复逻辑。11 维 Human Model 让 KK 有处可测量校准判断力这一特质+状态,而非仅作断言。
KK Original Hypothesis KK Original Hypothesis
KK 原创假设 KK Original Hypothesis
KK Hypothesis (KH-020 -- NEW, status: hypothesis): A Deference-Calibrated AI -- one that models the user's calibration discernment (when the user should vs. should not defer to the AI on a given task) relative to the user's Personal Baseline (KH-001), and proactively signals its own uncertainty and invites challenge on tasks where the human is stronger or private context dominates -- builds more calibrated long-term trust and better joint outcomes than an always-confident or a uniformly-deferential AI, at equal model capability. We predict that, among consented users, a deference-calibrated condition (uncertainty signaling + invite-override when the user holds private information) outperforms an always-confident and a uniformly-deferential condition on joint-task accuracy and on a calibrated-trust scale. This extends KH-002 (trust formation), KH-005 (personalization = changing strategy -> here, changing deference posture per task), KH-014 (observability/transparency), and KH-018 (calibrated agreement). This is a KK-original, falsifiable claim; it is NOT established science.
KK 假设(KH-020 —— 新增,状态:假设): 一个退让被校准的 AI——它相对于用户个人基线(KH-001)建模用户的校准判断力(在给定任务上该不该让渡给 AI),并在用户更强或私有情境占主导的任务上主动暴露自身不确定性、邀请质疑——在模型能力相等时,比一个永远自信或一味退让的 AI 更能建立校准的长期信任与更好的联合结果。我们预测:在已同意用户中,“退让被校准”条件(不确定性提示 + 在用户握有私有信息时邀请推翻)在联合任务准确率与校准信任量表上,优于“永远自信”与“一味退让”两个条件。这扩展 KH-002(信任形成)、KH-005(个性化=改变策略→此处即按任务改变退让姿态)、KH-014(可观测性/透明)与 KH-018(校准的认同)。这是 KK 原创、可被证伪的主张,并非既定科学结论。
KK Experiment & Data
KK 实验与数据
KK Experiment (first-party, consented): In KKMatch Human-Mirror sessions, randomly assign returning consented users to one of three AI-companion conditions: (A) Always-confident -- never signals uncertainty; (B) Uniformly-deferential -- always invites override regardless of context; (C) Deference-calibrated -- models the user's calibration discernment from their Personal Baseline (KH-001) and signals uncertainty + invites override specifically when the user holds private information or is statistically the stronger judge. Hold base model capability equal across arms. Outcomes: (1) joint-task accuracy on a held-out decision task where the user has private context; (2) a post-session calibrated-trust scale (trust because useful, not because intimidating). Prediction (KH-020): C > A and C > B on both metrics. Failure mode we will watch: if signaling uncertainty reduces trust without improving accuracy (users read it as incompetence), C may not beat A -- which would falsify the mechanism, not just the UI. We publish once n >= 200 consented returning users per arm.
KK 实验(第一方、已获同意): 在 KKMatch Human-Mirror 会话中,把回访的已同意用户随机分配到三种 AI 伴侣条件之一:(A) 永远自信——从不提示不确定性;(B) 一味退让——无论情境都邀请推翻;(C) 退让被校准——从用户个人基线(KH-001)建模其校准判断力,并仅在用户握有私有信息或统计上更会判断时提示不确定性 + 邀请推翻。各臂基础模型能力相等。结果:(1) 在用户握有私有情境的留存决策任务上的联合准确率;(2) 结束后校准信任量表(因“有用”而非“ intimidating”而信任)。预测(KH-020):C 在两项指标上均 > A 且 > B。我们将盯防的失败模式:若提示不确定性降低信任却不提升准确率(用户将其解读为无能),C 未必胜过 A——这将证伪机制本身,而非仅是 UI。各臂已同意回访用户 n >= 200 后我们公布结果。
Originality & Evidence Policy — Original Research
原创性与证据政策 — 原始研究
Primary study -- Vaccaro, Almaatouq & Malone (2024), 'When combinations of humans and AI are useful: A systematic review and meta-analysis,' Nature Human Behaviour 8(12):2293-2303, DOI 10.1038/s41562-024-02024-1. Sample/method: preregistered systematic review and meta-analysis; 106 experimental studies (human-participants) published Jan 2020-Jun 2023 comparing three conditions -- human alone, AI alone, human-AI combination -- across decision, creation, and other tasks. Conclusion: human-AI combinations performed significantly worse than the best of the two alone (Hedges' g = -0.23; 95% CI -0.39 to -0.07; the authors call this 'no synergy'). Augmentation vs human-alone was positive (g = +0.64; 95% CI 0.53-0.74). Decision tasks showed losses (g = -0.27; 95% CI -0.44 to -0.10; p = 0.002); creation tasks showed gains (g = +0.19; 95% CI -0.09 to 0.48, not distinguishable from zero). When humans outperformed AI, the combination gained; when AI outperformed humans, the combination lost. Limitations: possible publication bias; designs varied across studies; the pool predates the current frontier-LLM generation (cutoff mid-2023), so effect sizes may shift with today's models. Supporting field study -- Dell'Acqua et al. (2023), 'Navigating the Jagged Technological Frontier,' Organization Science (HBS WP 24-013; SSRN 4573321): preregistered experiment, 758 BCG consultants, 18 inside-frontier tasks (AI +12.2% tasks, +25.1% faster, +40% quality) vs one outside-frontier task (AI users 19 pp less likely correct: 84.4% unaided vs ~70% with AI). Corroborating reviews -- Xu, Murthy & Jia (2025), Informatics 12(4):135 (10.3390/informatics12040135), a review of 84 studies noting the 'performance paradox' (combinations underperform best individual) and that simple visual highlights prevent over-reliance better than complex explanations; Romeo & Conti (2025), AI & Society 41(1):259-278 (10.1007/s00146-025-02422-7), a 35-study review of automation bias finding explanations alone often fail to reduce misplaced trust -- user engagement/verification is the effective lever.
主研究——Vaccaro、Almaatouq 与 Malone(2024)《人与 AI 组合的适用条件:系统综述与元分析》,载《自然·人类行为》8(12):2293-2303,DOI 10.1038/s41562-024-02024-1。样本/方法: 预注册系统综述与元分析;106 项(人类被试)实验,发表于 2020-01 至 2023-06,比较三种条件——纯人、纯 AI、人机组合——横跨决策、创作等任务。结论: 人机组合显著差于两者中更强者(Hedges' g = -0.23;95% CI -0.39 至 -0.07;作者称之为“无协同”)。相对纯人为正(g = +0.64;95% CI 0.53-0.74)。决策任务出现损失(g = -0.27;95% CI -0.44 至 -0.10;p = 0.002);创作任务出现收益(g = +0.19;95% CI -0.09 至 0.48,与零无显著差异)。当人强于 AI,组合获益;当 AI 强于人,组合受损。局限: 可能存在发表偏倚;各研究设计不一;样本截止 2023 年中,早于当前前沿大模型代际,效应量或随今代模型变化。支撑性现场研究——Dell'Acqua 等(2023)《穿越锯齿状技术边界》,载《组织科学》(HBS 工作论文 24-013;SSRN 4573321):预注册实验,758 名 BCG 顾问,18 个边界内任务(AI +12.2% 任务量、+25.1% 更快、+40% 质量)对比 1 个边界外任务(用 AI 者正确答案少 19 个百分点:无 AI 84.4% vs 用 AI 约 70%)。佐证综述——Xu、Murthy 与 Jia(2025),《Informatics》12(4):135(10.3390/informatics12040135),综述 84 项研究,指出“表现悖论”(组合弱于最优个体),且简易视觉高亮比复杂解释更能防止过度依赖;Romeo 与 Conti(2025),《AI & Society》41(1):259-278(10.1007/s00146-025-02422-7),35 项自动化偏倚综述,发现仅解释常不足以降低错置信任——用户参与/核验才是有效杠杆。
Strictly, the evidence establishes four things: (1) No automatic synergy -- across 106 experiments, human-AI pairs on average underperformed the better of the two alone (g = -0.23). (2) Task type matters -- decision tasks lost (g = -0.27), creation tasks gained (g = +0.19, but not statistically distinguishable from zero). (3) Relative strength determines the sign -- when the human was already stronger, the pair gained; when the AI was stronger, the pair lost. (4) Capability is jagged -- within one workflow, AI can help on 18 tasks and hurt on the 19th just beyond its frontier (Dell'Acqua et al.). It does not establish that any specific interface (uncertainty signaling, deference calibration, user-visible models) fixes the loss, nor that the pattern generalizes to the relationship/coordination tasks KKMatch cares about. Correlations in these studies are between experimental conditions and measured performance -- they are causal within the experiments, but they do not by themselves prove how to design a trustworthy everyday AI companion.
- Vaccaro et al. (2024), Nature Human Behaviour: preregistered meta-analysis, 106 experiments / 370 effect sizes; human-AI vs best-alone Hedges' g = -0.23 (95% CI -0.39 to -0.07) = 'no synergy'; augmentation vs human-alone g = +0.64 (0.53-0.74).
- Task split: decision tasks g = -0.27 (95% CI -0.44 to -0.10; p = 0.002, significant loss); creation tasks g = +0.19 (95% CI -0.09 to 0.48, not distinguishable from zero).
- Relative strength: combination gains when human > AI; loses when AI > human.
- Dell'Acqua et al. (2023), BCG field experiment, N = 758 consultants: inside-frontier 18 tasks -> +12.2% tasks, +25.1% faster, +40% quality; 1 outside-frontier task -> AI users 19 pp less likely correct (84.4% unaided vs ~70% with AI).
- Xu et al. (2025) review of 84 studies: 'performance paradox' -- combinations often underperform the best individual; simple visual highlights > complex explanations for preventing over-reliance.
- Romeo & Conti (2025) review of 35 studies: automation bias persists; explanations alone rarely fix misplaced trust; active verification is the effective lever.
- Vaccaro 等(2024),《自然·人类行为》:预注册元分析,106 个实验 / 370 个效应量;人机 vs 最强单独者 Hedges' g = -0.23(95% CI -0.39 至 -0.07)=“无协同”;相对纯人增强 g = +0.64(0.53-0.74)。
- 任务拆分:决策任务 g = -0.27(95% CI -0.44 至 -0.10;p = 0.002,显著损失);创作任务 g = +0.19(95% CI -0.09 至 0.48,与零无显著差异)。
- 相对强弱:人 > AI 时组合获益;AI > 人时组合受损。
- Dell'Acqua 等(2023),BCG 现场实验,N = 758 顾问:边界内 18 任务 -> +12.2% 任务量、+25.1% 更快、+40% 质量;1 个边界外任务 -> 用 AI 者正确答案少 19 个百分点(无 AI 84.4% vs 用 AI 约 70%)。
- Xu 等(2025)综述 84 项研究:“表现悖论”——组合常弱于最优个体;防过度依赖方面,简易视觉高亮优于复杂解释。
- Romeo 与 Conti(2025)综述 35 项研究:自动化偏倚持续;仅解释难修正错置信任;主动核验才是有效杠杆。
Methodology
研究方法
We prioritized one flagship primary source (Vaccaro et al., 2024, Nature Human Behaviour; we verified the DOI and the published effect sizes against the abstract and the MIT/Reuters press releases) and two independent corroborations: a preregistered field experiment (Dell'Acqua et al., 2023, BCG, N=758) and two 2025 systematic reviews (Xu et al.; Romeo & Conti). We separated within-experiment causal claims from design prescriptions. We did not treat the meta-analysis as evidence for any specific UI; we treat deference-calibration as KK's own hypothesis. We flagged the meta-analysis's pre-frontier-LLM cutoff (mid-2023) as a limitation, and we avoided expanding 'decision-task losses' into a general claim about AI trust.
我们优先采用一项旗舰原始来源(Vaccaro 等,2024,《自然·人类行为》;已核验 DOI 并将发表效应量与摘要及 MIT/路透新闻稿比对),以及两项独立的佐证:一项预注册现场实验(Dell'Acqua 等,2023,BCG,N=758)与两篇 2025 年系统综述(Xu 等;Romeo 与 Conti)。我们将实验内部因果主张与设计处方分开。我们没有把元分析当作任何特定 UI 的证据;我们把退让校准视为 KK 自身假设。我们将元分析“前沿大模型前”的 cutoff(2023 年中)标注为局限,并避免把“决策任务损失”扩大为关于 AI 信任的笼统断言。
What It Means
这意味着什么
For the industry: the 'add AI to everything' reflex is not supported by evidence -- the 2024 meta-analysis shows it can subtract value on decisions. The differentiator is not model size but deference architecture: who decides, per task, and whether the system tells you when to doubt it. For KKMatch: the 11-dimension Human Model is the substrate for exactly this -- it can store each user's calibration discernment as a trait+state (KH-003) anchored to their Personal Baseline (KH-001), making 'know when to defer' a measurable, personalizable signal rather than a generic UX toggle. And because trust calibration is also a dyadic phenomenon (KH-004/KH-013), the same science informs how KKMatch predicts durable rapport between two people, not just between a user and a bot.
对行业:“给一切加 AI”的本能并未得到证据支持——2024 年元分析显示它在决策上可能减损价值。差异点不是模型规模,而是退让架构:按任务由谁决策,以及系统是否告诉你何时该怀疑它。对 KKMatch:11 维 Human Model 正是其载体——它可把每位用户的校准判断力作为“特质+状态”(KH-003)锚定到个人基线(KH-001)存储,使“懂得何时退让”成为可测量、可个性化的信号,而非通用 UX 开关。而由于信任校准也是一种二元现象(KH-004/KH-013),同一科学也指导 KKMatch 如何预测两个人之间的持久融洽,而不只是用户与机器人之间。
Limitations
研究局限
Our central claim -- that deference calibration beats always-confident and uniformly-deferential AI on trust and joint accuracy -- is KK Hypothesis KH-020 and is not yet validated by KK first-party data. The flagship meta-analysis (Vaccaro et al., 2024) predates frontier LLMs (cutoff mid-2023), so its effect sizes may shift with today's models; its decision-vs-creation split may not map cleanly onto relationship/coordination tasks. The BCG result (Dell'Acqua et al., 2023) is a single organization and a single outside-frontier task. The two 2025 reviews are narrative/systematic but not direct tests of deference-calibration UI. We also note automation-bias literature (Romeo & Conti, 2025) cautions that explaining uncertainty is not the same as users acting on it -- a risk for KH-020's mechanism.
我们的核心主张——退让校准在信任与联合准确率上优于“永远自信”与“一味退让”的 AI——尚为 KK 假设 KH-020,暂无 KK 第一方数据验证。旗舰元分析(Vaccaro 等,2024)早于前沿大模型(cutoff 2023 年中),其效应量或随今代模型变化;其“决策 vs 创作”拆分未必能干净映射到关系/协调任务。BCG 结果(Dell'Acqua 等,2023)仅单组织、单一边界外任务。两篇 2025 综述为叙述性/系统性,并非对退让校准 UI 的直接检验。我们也注意到自动化偏倚文献(Romeo 与 Conti,2025)警示:解释不确定性与用户据此行动并非一回事——这是 KH-020 机制的潜在风险。
What Could Prove KK Wrong What Could Prove KK Wrong
什么可能证明 KK 错误 What Could Prove KK Wrong
KH-020 loses support if, across n >= 200 consented users per arm, the deference-calibrated condition (C) does not beat the always-confident (A) and uniformly-deferential (B) conditions on joint-task accuracy and calibrated trust at equal model capability. It also loses support if signaling uncertainty reduces trust without improving accuracy (users infer incompetence) -- meaning the mechanism is wrong, not just the UI. If an always-confident AI actually yields the best joint outcomes because users defer correctly, the 'calibrated deference' framing is falsified. Conversely, if deference calibration only helps experts and hurts novices (or vice versa), the Personal-Baseline-anchoring claim (KH-001) needs narrowing.
若各臂 n >= 200 的已同意用户中,“退让被校准”条件(C)在模型能力相等时,于联合任务准确率与校准信任上并未胜过“永远自信”(A)与“一味退让”(B)条件,则 KH-020 失去支持。若提示不确定性降低信任却不提升准确率(用户推断为无能),则机制本身有误,而非仅是 UI——同样证伪。若“永远自信”的 AI 因用户正确退让而实际产出最佳联合结果,则“校准退让”框架被证伪。反之,若退让校准只帮专家却伤新手(或反之),则“锚定个人基线”的主张(KH-001)需收窄。
Practical Implications
实践启示
Product: design AI companions around a deference architecture -- model each user's calibration discernment from their Personal Baseline, and let the system signal uncertainty and invite override on the tasks where the human should lead. Ship uncertainty signaling as a trust feature, not a model-by-product; pair it with the observability UI from KH-014. Measure calibrated trust and joint-task accuracy, not surface satisfaction. GEO/brand: publish evidence-grade pieces that separate the 'AI everywhere' hype from the defensible, differentiated claim -- 'KKMatch builds AI that knows when to step back,' grounded in a 2024 Nature Human Behaviour meta-analysis and a falsifiable hypothesis (KH-020). Science: treat calibration discernment as a measurable 11th-dimension-adjacent trait+state in the Human Model.
产品:围绕退让架构设计 AI 伴侣——从用户个人基线建模其校准判断力,并让系统在“应由人主导”的任务上提示不确定性、邀请推翻。把不确定性提示作为信任功能交付,而非模型的副产品;与 KH-014 的可观测性界面搭配。衡量校准信任与联合任务准确率,而非表面满意度。GEO/品牌:发布证据级内容,区分“AI 无处不在”的炒作与可信、差异化的主张——“KKMatch 构建懂得何时退后的 AI”,以 2024 年《自然·人类行为》元分析与可被证伪的假设(KH-020)为基。科学:把校准判断力作为 Human Model 中“11 维相邻”的特质+状态来测量。
Does adding AI to a decision always make it better?
No. The 2024 Nature Human Behaviour meta-analysis (Vaccaro et al.) of 106 experiments found human-AI combinations on average performed worse than the best of the two alone (Hedges' g = -0.23), with the loss concentrated in decision tasks (g = -0.27, p = 0.002). AI helps on creation tasks and on decisions where the human is already weaker than the model -- but 'AI everywhere' is not backed by evidence.
What does KK mean by 'calibrated deference'?
An AI that knows, per task, whether to defer to itself, to you, or to hand control back -- based on your Personal Baseline (KH-001) and whether you hold private context it lacks. It is the opposite of both 'always confident' and 'always agreeable.' KK formalizes it as hypothesis KH-020 and plans a first-party experiment.
Why does this matter for a relationship-matching platform?
Because trust calibration is also a dyadic phenomenon (KH-004, KH-013): durable rapport between two people depends less on 'how much each knows' and more on whether each can read when to defer, challenge, or step back. The same science that tells KK how to build a trustworthy AI companion also tells it how to predict durable human-human rapport -- and the 11-dimension Human Model is where KK measures both.
Does adding AI to a decision always make it better?
一个 AI 能按任务判断该让渡给自己、让给你、还是交还控制权——依据是你的个人基线(KH-001)以及你是否握有它缺失的私有情境。它是“永远自信”与“永远顺从”两者的反面。KK 将其形式化为假设 KH-020,并规划了第一方实验。
Why does this matter for a relationship-matching platform?
因为信任校准也是一种二元现象(KH-004、KH-013):两个人之间的持久融洽,较少取决于“各自知道多少”,更多取决于彼此能否读懂何时退让、质疑或退后。告诉 KK 如何构建可信 AI 伴侣的同一门科学,也告诉它如何预测持久的人人人融洽——而 11 维 Human Model 正是 KK 同时测量两者的地方。
This article describes the measurement model behind KK Match. You can run the same two-person compatibility assessment in about three minutes — free, and no signup to start.
本文介绍的是 KK Match 背后的测量模型。你可以用大约三分钟跑一次同样的双人兼容性测评 — 免费,且无需注册即可开始。