KKResearch研究

KK Research / Effect Reports / Report 03

KK 研究 / 效果实测报告 / 报告 03

Signal Coverage 通道到位率

Text, voice, and face: how many of the three channels actually arrive?

文本、语音、画面:三路信号实际到了几路?

Across 21 first-party training samples, text is present in every one, voice prosody in 12, and face signal in exactly 1. Multimodal is, today, text plus half of voice — and we would rather publish that than advertise a capability we cannot deliver.

在 21 份第一方训练样本中,文本通道每份都到位,语音韵律 12 份,画面信号恰好 1 份。多模态在今天是「文本 + 半个语音」——我们宁愿公布这一点,也不去宣传做不到的能力。

KKMatch Human Intelligence Research Team KKMatch 人类智能研究团队 · Published: 2026-09-12发布于:2026-09-12 · 6 min read阅读约 6 分钟
21/21
Text channel文本通道
12/21
Voice prosody语音韵律
1/21
Face signal画面信号
29
Mic grants麦克风授权
1
Camera grants摄像头授权

The Challenge

挑战

Multimodal understanding is one of the four differentiators this product claims, and it is the easiest one to fake. A landing page can promise "reads your voice and your face" and no visitor will ever be able to check. Internal engineering has the opposite problem: it knows exactly which arrays are empty. This report closes that gap by asking a deliberately unglamorous question — for each signal the product implies it uses, how many stored samples actually contain it?
多模态理解是本产品宣称的四项差异化之一,而它恰恰是最容易造假的一项。落地页可以承诺「读取你的声音与表情」,而没有访客能验证。内部工程则面对相反的问题:它精确知道哪些数组是空的。本报告用一个刻意朴素的提问来弥合这个落差——对产品暗示会使用的每一路信号,实际有多少份落库样本真的包含它?
The rule we hold ourselves to: a capability is described on the site only at the coverage level this report can demonstrate. Where coverage is near zero, the site says the capability is not offered — rather than describing it in the present tense and letting users assume it works. 我们自我约束的规则: 站点上对一项能力的描述,只写到本报告能证明的覆盖水平。覆盖率接近零的能力,站点就明说「未提供」——而不是用现在时描述它、让用户默认它已经工作。

The Approach

做法

  1. Start from the stored samples. Every record in the training-sample store is examined — 21 of them — and each optional signal field is tested for presence and for non-empty content. A field holding an empty array counts as absent.从落库样本出发。 逐一检查训练样本库中的每条记录——共 21 条——对每个可选信号字段同时检验「字段存在」与「内容非空」。字段存在但数组为空,计为缺失。
  2. Count frames, not just fields. Where a signal is present we also record how many frames it carried, because a two-frame "prosody track" is technically present and practically useless.数字段,也数帧数。 对已到位的信号,另外记录其携带的帧数——因为一条两帧的「韵律轨」在技术意义上存在,在实用意义上无用。
  3. Cross-check against consent events. Capture cannot exceed permission: if the camera was authorised once, no more than one session can possibly hold face data. Comparing the two independently derived counts is how we detect a silent capture failure.与授权事件交叉核对。 采集量不能超过授权量:如果摄像头只被授权过一次,那么持有画面数据的会话不可能超过一场。把这两个独立来源的数字对照,是发现「静默采集失败」的方法。
  4. Report the funnel, not the peak. The headline figures are end-to-end arrival rates at the store. We do not report "the capture pipeline supports X" — only what survived all the way into a stored sample.报告漏斗,而非峰值。 核心数字是「最终抵达存储的端到端到位率」。我们不报告「采集管线支持 X」——只报告一路活到落库样本的量。

The Result

结果

The three channels are not three-quarters full, or half full. They are one full channel, one half channel, and one channel that has essentially never run.
三路通道不是「八分满」或「一半满」。它们是:一路满的、一路半的、以及一路基本从未跑起来的。
Channel通道 Samples with signal含信号的样本 Arrival到位率 Frame depth帧深
Text turns (baseline)文本轮次(基线) 21 / 21100% full transcript完整文稿
Voice prosody (energy / f0 / voiced)语音韵律(能量 / 基频 / 有声) 12 / 2157.1% 2 – 121 frames (mean 50)帧(均值 50)
Face signal (smile / blink / mouth / brow)画面信号(微笑 / 眨眼 / 口型 / 眉) 1 / 214.8% 6 frames
Turn timing (reply latency, lengths)轮次时序(回复延迟、长度) 12 / 2157.1% 35 turn pairs对轮次

The consent cross-check

授权交叉核对

Camera coverage is not merely low — it is low for a discoverable reason. Only one camera-permission event exists in the entire instrumented history, and exactly one sample holds face data. Permission and capture agree, so this is a demand-side absence, not a broken pipeline.
画面覆盖率低,而且低得原因可查。在全部埋点历史中,摄像头授权事件只出现 1 次,而持有画面数据的样本也恰好 1 份。授权与采集相互吻合,所以这是需求侧的空缺,不是管线故障。
Permission event授权事件 Times raised出现次数 Granted通过
Microphone麦克风2929
Camera摄像头11
Two further gaps are visible in the same pass. First, sessions declared as video mode outnumber samples containing any face data by 5 to 1 — the mode exists in the product before the signal does. Second, only 21 of 74 sessions produced a training sample at all (28.4%), so even the text channel is not fully harvested.
同一次核对还暴露两处缺口。其一,声明为视频模式的会话数,是含任何画面数据的样本数的 5 倍——在产品里「模式」先于「信号」存在。其二,74 场会话中只有 21 场产出了训练样本(28.4%),所以即便是文本通道也并未全量收割。

A note on turn timing

关于轮次时序的说明

The turn-timing signal records the gap between one turn ending and the next beginning. Median 7.5 s, mean 17.6 s, 90th percentile 48.5 s, maximum 49.7 s. This is not system latency: it includes the user's own thinking and typing time, and no attempt has been made to separate the two. We label it as a turn gap rather than a response time so that it cannot be quoted as a latency benchmark.
轮次时序信号记录的是「上一轮结束」到「下一轮开始」之间的间隔。中位数 7.5 秒,均值 17.6 秒,90 分位 48.5 秒,最大 49.7 秒。这不是系统延迟: 其中包含用户自己的思考与打字时间,且我们没有尝试把两者分离。我们把它标注为「轮次间隔」而非「响应时间」,以免它被引用为时延基准。

What It Means

含义

The honest framing: one channel is complete, one is half-complete, one is unused. Stating that precisely is worth more than a plausible-sounding multimodal claim, because it tells a prospective user, a partner and our own roadmap exactly where the next unit of work pays off. 诚实的表述: 一路通道完整,一路半完整,一路未启用。精确说明这一点,比一句听起来合理的多模态宣称更有价值——因为它同时告诉潜在用户、合作方和我们自己的路线图:下一份投入应该花在哪里。

Reproduce it

复现方式

Pure aggregation over two stores — no probe required. For each record in data/mirror_training_samples.jsonl, test whether voice_prosody, video_signal_seq and turn_timing are present and non-empty. Permission counts come from data/mirror_analytics.json, filtered on event_type equal to mic_permission and camera_permission. 对两个存储的纯聚合,无需探针。对 data/mirror_training_samples.jsonl 中的每条记录,检验 voice_prosodyvideo_signal_seqturn_timing 是否存在且非空。授权计数来自 data/mirror_analytics.json,按 event_type 等于 mic_permissioncamera_permission 过滤。

Limitations

研究局限

What would prove this report wrong Falsifiable

什么能推翻本报告 可证伪

Back to Effect Reports返回效果实测报告