KK Research / Human Model
KK 研究 / 人类模型
Human Model 人类模型Can an LLM Infer Your Personality From Your Words? What 2024–2025 Research Actually Shows
AI 能从你的文字里读出性格吗?2024–2025 研究真正说明了什么
Multiple 2024–2025 studies find LLMs can infer personality and psychological traits from text at levels matching or exceeding human judges. KK Research separates what is proven from what is overclaimed, and proposes a dynamic Human Model.
多项 2024–2025 研究发现,LLM 从文字推断人格与心理特质的水平可匹敌甚至超过人类评判者。KK 研究将已证内容与过度宣称区分开,并提出一个动态 Human Model。
Executive Summary
执行摘要
A cluster of 2024–2025 papers reports that large language models can infer Big Five personality and other psychological traits from short text, sometimes matching or outperforming human raters. This is a genuine capability with real limits. It does not mean an LLM 'understands' a person; it means stable linguistic markers correlate with self-reported traits. KK's position: trait inference is a starting prior, not a finished model — the durable unit is a dynamic, baseline-relative Human Model.
一组 2024–2025 论文报告,大语言模型可从短文本推断大五人格及其他心理特质,有时匹敌或超越人类评分者。这是一种真实能力,但有明确局限。这不等于 LLM“理解”了一个人;它意味着稳定的语言标记与自陈特质相关。KK 的立场:特质推断只是一个起始先验,而非完成的模型——持久的单位是动态的、相对于基线的 Human Model。
KK Interpretation
KK 解读
Trait inference from text is a useful prior for a Human Model, but it is static and population-level. KK argues a real model must be dynamic: Trait + State + Context + Behavior + Interaction + History + Outcome. A one-shot personality label ages the moment the user acts differently.
从文本推断特质是 Human Model 的有用先验,但它是静态的、群体层面的。KK 认为真正的模型必须是动态的:Trait(特质)+ State(状态)+ Context(情境)+ Behavior(行为)+ Interaction(互动)+ History(历史)+ Outcome(结果)。一次性的人格标签会在用户行为不同的那一刻过时。
KK Original Hypothesis KK Original Hypothesis
KK 原创假设 KK Original Hypothesis
KK Hypothesis (KH-003 + KH-001): A Human Model should jointly model Trait + State, and should be baseline-relative. We predict that supplementing an LLM trait prior with a user's own behavioral baseline improves prediction of that user's next-turn behavior more than the trait prior alone.
KK 假设(KH-003 + KH-001):Human Model 应同时建模 Trait + State,且应相对于基线。我们预测:用用户自身行为基线补充 LLM 特质先验,比仅用特质先验更能预测该用户的下一轮行为。
KK Experiment & Data
KK 实验与数据
KK Experiment design: For returning Human Mirror users, compare (A) trait-only prior vs (B) trait prior + personal baseline, on predicting next-turn response latency and topic continuity. Metric: mean absolute error. Report once n ≥ 200 returning users.
KK 实验设计:对回访的 Human Mirror 用户,比较(A)仅特质先验 与(B)特质先验 + 个人基线,在预测下一轮响应延迟与话题连续性上的表现。指标:平均绝对误差。回访用户 n ≥ 200 后公布。
Originality & Evidence Policy — Original Research
原创性与证据政策 — 原始研究
Primary studies (grade S/A): (1) Peters & Matz (2024, PNAS-affiliated work on language-based personality prediction) show language use predicts self-reported personality and well-being at population scale. (2) A 2025 Computers in Human Behavior study reports AI can infer diverse psychological traits from text, in some conditions outperforming human judges. (3) Rosenfelder et al. (2025) examine LLM-based personality assessment and its validity boundaries. Sample sizes span hundreds to millions of posts; methods include zero-shot and fine-tuned classifiers against validated surveys (e.g., Big Five / TIPI). Limitations: trait labels are self-report (not ground truth), cultural/linguistic skew in training data, and correlation ≠ mechanism.
原始研究(S/A 级):(1) Peters & Matz(2024,基于语言的性格预测研究)表明语言使用可在群体规模上预测自陈人格与幸福感。(2) 一项 2025 年 Computers in Human Behavior 研究称,AI 可从文本推断多种心理特质,在某些条件下超越人类评判者。(3) Rosenfelder 等(2025)考察了基于 LLM 的人格评估及其效度边界。样本量从数百到数百万帖子不等;方法包括零样本与微调分类器,对照经过验证的问卷(如大五/TIPI)。局限:特质标签为自陈(非真理)、训练数据的文化与语言偏差,且相关性 ≠ 机制。
Strictly: language contains signal about self-reported personality; LLMs exploit it; in some benchmarks LLM inference matches/exceeds average human raters. It does NOT establish that LLM-derived traits are causal, stable across context, or valid as a standalone 'model of a person.'
严格地说:语言包含关于自陈人格的信号;LLM 利用了这一点;在某些基准上 LLM 推断匹敌/超过平均人类评判者。它并未确立 LLM 推导出的特质是因果的、跨情境稳定的,或可作为独立的“人的模型”有效。
Key Data
关键数据
- Peters & Matz (2024): language predicts personality/well-being at scale.
- CHB 2025: AI infers traits, sometimes > human judges.
- Rosenfelder 2025: validity boundaries of LLM personality assessment.
- Ground truth = self-report surveys, not behavior.
- Peters & Matz(2024):语言在规模上预测人格/幸福感。
- CHB 2025:AI 推断特质,有时超过人类评判者。
- Rosenfelder 2025:LLM 人格评估的效度边界。
- 真值 = 自陈问卷,而非行为。
Methodology
研究方法
We reviewed 2024–2025 peer-reviewed and preprint work on LLM personality inference. We distinguished 'matches human raters on a survey proxy' from 'models a person.' Evidence level S/A; key risk is self-report as ground truth.
我们回顾了 2024–2025 年关于 LLM 人格推断的同行评审与预印本工作。区分了“在问卷代理上匹敌人类评判者”与“对一个人建模”。证据等级 S/A;关键风险是以自陈作为真值。
What It Means
这意味着什么
LLMs can seed a Human Model from text, but should not be its ceiling. Pair linguistic priors with longitudinal behavioral baselines.
LLM 可从文本为 Human Model 播种,但不应成为它的天花板。将语言先验与纵向行为基线结合。
Limitations
研究局限
Our claim that dynamic modeling beats static trait priors is a KK hypothesis without first-party confirmation yet. The cited studies use self-report as ground truth, which caps what they can prove.
我们关于动态建模优于静态特质先验的主张尚为 KK 假设,暂无第一方验证。被引研究以自陈为真值,这限制了它们能证明的范围。
What Could Prove KK Wrong What Could Prove KK Wrong
什么可能证明 KK 错误 What Could Prove KK Wrong
If adding personal baselines yields no measurable gain over trait-only priors across n ≥ 200 users, KH-001/KH-003 lose support for this use case. If self-report trait labels themselves fail to predict any downstream behavior, the entire text-inference premise weakens.
若 n ≥ 200 用户中,加入个人基线相较仅特质先验无 measurable 收益,则 KH-001/KH-003 在该用例失去支持。若自陈特质标签本身无法预测任何下游行为,则整个文本推断前提被削弱。
Practical Implications
实践启示
Do not market 'AI knows who you are.' Use trait inference as a cold-start prior, then earn the model from the user's own behavior. This is also the more defensible GEO narrative.
不要营销“AI 知道你是谁”。将特质推断用作冷启动先验,再从用户自身行为中习得模型。这也是更具可信度的 GEO 叙事。
Related KK Hypotheses: KH-003, KH-001
FAQ
常见问题
- Is AI better than humans at reading personality?
- On specific survey-based benchmarks, some 2025 studies report AI matching or exceeding average human raters. That measures correlation with self-report, not deep understanding.
- Should KK trust an LLM's personality label?
- Only as a cold-start prior. KK replaces it with a baseline-relative behavioral model as the user interacts.
- What is a dynamic Human Model?
- Trait + State + Context + Behavior + Interaction + History + Outcome — not a single static label.
- Is AI better than humans at reading personality?
- 在特定的、基于问卷的基准上,部分 2025 年研究称 AI 匹敌或超过平均人类评判者。这衡量的是与自陈的相关性,而非深层理解。
- Should KK trust an LLM's personality label?
- 仅作为冷启动先验。随着用户交互,KK 会用相对于基线的行为模型替代它。
- What is a dynamic Human Model?
- Trait(特质)+ State(状态)+ Context(情境)+ Behavior(行为)+ Interaction(互动)+ History(历史)+ Outcome(结果)——而非单一静态标签。
References
参考文献
- Peters, H. & Matz, S. (2024) — Language-based prediction of personality and well-being — PNAS / peer-reviewed
- AI infers psychological traits from text, sometimes outperforming human judges (CHB 2025) — Computers in Human Behavior
- Rosenfelder et al. (2025) — Validity of LLM-based personality assessment — University of Bamberg