KKResearch研究

KK Research / Effect Reports / Report 02

KK 研究 / 效果实测报告 / 报告 02

Honest Confidence 置信诚实性

74 real conversations, 35 models, and a 40% fallback rate

74 场真实对话、35 份模型,以及 40% 的降级率

Every one of the 385 trait slots was filled — yet 40% of the models returned nothing except an explicit "insufficient conversation data" flag, and only 17% reached an overall confidence of 0.6. Trait coverage is not trait quality.

385 个特质槽位全部被填充——然而 40% 的模型除了显式的「对话数据不足」标记外什么都没给出,仅 17% 达到 0.6 的整体置信度。特质覆盖率不等于特质质量。

KKMatch Human Intelligence Research Team KKMatch 人类智能研究团队 · Published: 2026-09-12发布于:2026-09-12 · 7 min read阅读约 7 分钟
74
Real sessions真实会话
35
Models produced产出的模型
100%
Trait slot coverage特质槽位覆盖率
0.400
Median confidence置信度中位数
40.0%
Explicit fallback rate显式降级率

The Challenge

挑战

An engine that always returns a full profile looks more impressive than one that sometimes says "I don't know". It is also more dangerous. Real conversations are wildly unequal in length: in our data the median session is 3 turns, and 47.3% are a single turn — a greeting, or a lost visitor. An engine optimised to always produce output will happily generate eleven confident-looking numbers from one sentence.
一个永远返回完整画像的引擎,看起来比一个有时会说「我不知道」的引擎更体面。但它也更危险。真实对话的长度极不均衡:在我们的数据里,会话轮数中位数是 3 轮47.3% 只有 1 轮——一句问候,或一个走失的访客。一个以「永远产出」为目标的引擎,会乐于从一句话里生成十一个看起来很有把握的数字。
The question this report answers is therefore not "how good are the scores" but "does the engine say so when it does not have enough to go on?" That is the load-bearing claim behind not fabricating user profiles.
因此本报告要回答的不是「分数有多好」,而是「在信息不足时,引擎会不会主动说出来?」 这正是「不编造用户画像」这一承诺的承重墙。

The Approach

做法

  1. Take the full real corpus. Every session in the live store is included — 74 of them, across text, voice and video modes. No record is filtered out for being too short, which is exactly the subset that would flatter the numbers.取全量真实语料。 线上存储中的每场会话都纳入——共 74 场,覆盖文本、语音、视频三种模式。不因「太短」剔除任何记录——而恰恰是这部分记录会对指标最有利。
  2. Read the generated models, not the events. Confidence figures are taken from the model records themselves, because the model store is authoritative while the client event stream is best-effort instrumentation.读生成的模型,而非事件。 置信度数字取自模型记录本身——模型存储是权威来源,而客户端事件流只是尽力而为的埋点。
  3. Count slots, not sessions. Each model has 11 dimensions, giving 35 × 11 = 385 slots. Coverage and confidence are reported at slot level so that a single missing dimension cannot hide inside a "model-level" average.数槽位,而不是数会话。 每份模型 11 维,共 35 × 11 = 385 个槽位。覆盖率与置信度按槽位统计,这样单个缺失维度无法藏在「模型级」平均值里。
  4. Separate "filled" from "meaningful". A slot counts as covered if it holds any value; it counts as earned only if the engine produced real evidence rather than its fallback text. The gap between the two is the headline of this report.把「填了」与「有内容」分开。 槽位只要有任何取值就算已覆盖;只有引擎给出真实证据、而非兜底文本时,才算已赢得。两者之间的缺口,就是本报告的核心结论。

The Result

结果

Coverage is perfect; usefulness is not. All 385 slots across all 35 models were filled — a 100% coverage rate. But the distribution of confidence tells a very different story.
覆盖率满分;可用性不是。 35 份模型的全部 385 个槽位都被填充——覆盖率 100%。但置信度的分布讲了另一个故事。
Reading读数 Value数值 Share占比
Models whose every dimension says only "insufficient conversation data"所有维度都只返回「对话数据不足」的模型14 / 3540.0%
Models whose every dimension sits at the neutral 50所有维度都停在中性值 50 的模型15 / 3542.9%
Models at overall confidence ≥ 0.6整体置信度 ≥ 0.6 的模型6 / 3517.1%
Models at overall confidence ≤ 0.35整体置信度 ≤ 0.35 的模型11 / 3531.4%
Trait slots sitting at exactly 0.40 confidence — the fallback bucket置信度恰好为 0.40 的特质槽位——兜底档153 / 38539.7%
Sessions completed (of 74)完成的会话(共 74 场)34 / 7445.9%
Overall confidence across the 35 models: median 0.400, mean 0.397, 90th percentile 0.650, maximum 0.800.
35 份模型的整体置信度:中位数 0.400,均值 0.397,90 分位 0.650,最高 0.800。

Confidence is not uniform across dimensions

置信度在各维度间并不均匀

Where the engine does commit, it commits unevenly. Observable, text-carried dimensions earn roughly 50% more confidence than the ones that depend on inference about a person's inner state.
当引擎确实给出判断时,它的把握也是不均匀的。可观察的、由文字直接承载的维度,比依赖内在状态推断的维度所赢得的置信度高出约五成。
Dimension维度 Mean conf.平均置信 σσ Relative相对水平
communication_style0.4600.225
directness0.4570.232
openness0.4500.194
engagement0.4390.199
ai_interaction_style0.4370.199
trust_disclosure_pattern0.4100.171
emotional_expression0.3930.155
reflection_style0.3760.196
decision_style0.3470.231
conflict_style0.3390.243
uncertainty_response0.3340.173
The pattern is coherent rather than random: communication_style and directness — things a reader can actually see in a transcript — lead at 0.46, while uncertainty_response and conflict_style — which require seeing a person under pressure, something a three-turn chat rarely supplies — trail at 0.33.
这个分布是有内在逻辑的,而非随机:communication_styledirectness——在文稿中真能看得见的东西——以 0.46 领先;而 uncertainty_responseconflict_style——需要观察一个人处在压力下的表现,而三轮回合很少提供这种机会——以 0.33 垫底。

Context: how sessions actually reach the engine

背景:会话如何抵达引擎

74 landing views落地页浏览
73 starts开始
36 consent granted授权通过
34 mirror started进入对话
34 sessions completed会话完成
This funnel is taken from the client event stream, which is best-effort instrumentation: session-completion events (35) and completed-session records (34) differ by one, and consent events (36) exceed mirror-start events (34). We publish the raw stream rather than silently reconciling it, and we do not use these counts as headline figures.
该漏斗取自客户端事件流,属尽力而为的埋点:会话完成事件(35)与已完成会话记录(34)相差 1,授权事件(36)多于进入对话事件(34)。我们公布原始事件流,而不是悄悄抹平差异,也不把这些计数当作主结论。

What It Means

含义

The 40% fallback rate is the most important number we have published, and it is not a good one. Here is what we take from it.
40% 的降级率是我们公布过的最重要的数字,而且它并不好看。我们的解读如下。
The honest framing: if we reported only the 100% coverage figure, this engine would look finished. Reporting the 40% fallback rate instead shows a measurement system that is calibrated to admit its own ignorance — which is the only state in which its confident readings deserve to be believed. 诚实的表述: 如果我们只公布 100% 覆盖率,这台引擎会显得已经完工。改为公布 40% 降级率,展现的是一套被标定为「肯承认自己无知」的测量系统——而只有在这种状态下,它那些有把握的读数才值得被相信。

Reproduce it

复现方式

Pure aggregation over the JSON stores — no probe required. Count models in data/mirror_models.json, sum their traits[].confidence values, and count sessions in data/mirror_sessions.json. The fallback rate is the share of models whose every trait carries evidence text beginning "Insufficient". 对 JSON 存储的纯聚合,无需探针。统计 data/mirror_models.json 中的模型数,汇总其 traits[].confidence;统计 data/mirror_sessions.json 中的会话数。降级率 = 所有特质的证据文本均以 "Insufficient" 开头的模型占比。

Limitations

研究局限

What would prove this report wrong Falsifiable

什么能推翻本报告 可证伪

Back to Effect Reports返回效果实测报告