跳到主内容
@wquguru
精选75Rohan Paul论文研究

措辞变化可让LLM接受虚假信息

LLMs can accept the same false claim differently depending on its tone, certaint…

原文
发到 X

LLMs can accept the same false claim differently depending on its tone, certainty, and grammatical form.

Small wording changes can make LLMs accept false claims, while larger and instruction-tuned models resist them more.

Models must decide whether to trust a user’s new claim or rely on facts stored during training.

EoBench tests this choice with about 66K false claims written in 19 styles across form, evidence, certainty, and tone.

The team evaluated 18 Gemma, Llama, and Qwen models, then kept cases where each model already knew the correct fact.

Commands, child-directed wording, formal language, and authority claims persuaded models most, while weak claims and counterfactuals persuaded them least.

Across Llama and Gemma, larger models followed false context less often, and instruction tuning usually reduced that behavior.

The finding shows that prompt wording can quietly change model answers, so evaluations and product safeguards must test linguistic framing directly.

– arxiv. org/abs/2607.18232

Title: "It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief"

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近