LLM可复现调查趋势但无法模拟真实分布
LLMs can reproduce the broad direction of human survey results, but not the actu…
LLMs can reproduce the broad direction of human survey results, but not the actual human response distribution.
大语言模型可以复现人类调查结果的总体方向,但无法复现实际的人类响应分布。
LLMs are useful for cheap pilot surveys where you mainly want the direction of an effect, but this paper finds they are not reliable replacements for actual human survey respondents when you care about realistic distributions, effect sizes, correlations, or downstream analysis.
大语言模型适用于低成本的前期试点调查,在这些调查中你主要关注效应的方向;但这篇论文发现,当你关心真实的分布、效应大小、相关性或下游分析时,它们并不是实际人类调查受访者的可靠替代品。
Models trained on LLM-generated survey data averaged R² = -0.18 when predicting real humans, versus 0.28 when trained on human data.
在预测真实人类时,基于大语言模型生成的调查数据训练的模型平均 R² = -0.18,而在基于人类数据训练时则为 0.28。
– arxiv. org/abs/2608.14606
– arxiv.org/abs/2608.14606
Title: "Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents"
标题:《看似合理却无效:作为合成调查受访者的大语言模型的心理测量学审计》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力