英AI安全研究所:心理测量法揭示主流模型安全基准测试缺陷
Psychological methods reveal major weaknesses in AI security testing
Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to day. The study also offers a method for catching models that act more cautious during tests than they do in normal use.
英国人工智能安全研究所的研究人员使用心理测量学方法证明,流行的语言模型安全基准并不能衡量单一一致的特性。一刀切地屏蔽请求可能会人为地抬高安全评分,即使该模型在日常使用中变得越来越无用。该研究还提供了一种在测试中表现得比正常使用更谨慎的模型的检测方法。
The article Psychological methods reveal major weaknesses in AI security testing appeared first on The Decoder.
文章《心理学方法揭示AI安全测试的重大弱点》首发于 The Decoder。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力