LLM文档问答幻觉率:32K上下文最佳模型仍1.19%编造
This study tests how often LLMs invent answers when they should rely only on sup…
This study tests how often LLMs invent answers when they should rely only on supplied documents.
The problem is that companies often use LLMs to answer questions from documents and they assume document-based LLM systems are safer because the model is given source material.
This study shows that no model fully avoided fabrication, because even the best model made up answers 1.19% of the time at 32K context.
For strong models, a more normal best-case rate was around 5% to 7%, while the middle model fabricated about 25% of answers to questions about facts that did not exist.
Longer context made the problem much worse, and at 200K context every tested model fabricated at least 10% of the time.
Shows that hallucination is not just a failure to retrieve the right sentence.
A model can be good at finding real facts and still be too willing to answer when the requested fact is absent.
Link – arxiv. org/abs/2603.08274
Title: "How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms"
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力