最强LLM也无法完全免疫越狱攻击
Perfect immunity from jailbreak is not possible even for the strongest of LLMs.
Perfect immunity from jailbreak is not possible even for the strongest of LLMs.
New study shows that frontier models are getting harder to jailbreak, but not impossible to jailbreak.
The study attacks Anthropic’s Fable 5 and Opus 4.8 with automated red-team tools that keep rewriting harmful prompts until the model either refuses or gives a bad answer.
Fable 5 was more robust than Opus 4.8, with its worst attack success rate at 6.1%, while Opus 4.8 reached 11.5% under the strongest attack.
The hard truth is that avoiding absolutely every jailbreak is practically impossible, because even a tiny failure rate can produce many harmful completions when attacks are automated and repeated at scale.
The most crucial point is, that the old cartoon version of jailbreaks, weird encodings and theatrical role-play, is no longer the main problem.
The surviving weakness is contextual, because adaptive attackers rewrite the request after refusals, searching for a frame the model treats as legitimate rather than dangerous.
That is why perfect immunity is probably the wrong target; language models do not inspect intent from a clean moral altitude, they infer meaning through phrasing, context, and precedent.
In any system this flexible, there will always be boundary cases where a harmful request looks enough like education, safety research, fiction, troubleshooting, or policy analysis to slip through.
Link – arxiv. org/abs/2606.18193
Title: "A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models"
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力