The Growing Risk of “Multi-Turn Attacks” on Open-Weight Large Language Models
Large Language Models (LLMs) are rapidly transforming industries, offering unprecedented capabilities in automation, content creation, and data analysis. Though, a recent study by Cisco researchers highlights a critical vulnerability: open-weight llms are increasingly susceptible to complex attacks that bypass built-in safety measures over extended interactions. This poses meaningful operational and ethical risks for organizations deploying these powerful tools.
This article delves into the findings of the Cisco research, explaining what “multi-turn attacks” are, which models are most vulnerable, and what steps organizations can take to mitigate these emerging threats.
What are Multi-Turn Attacks and Why are They Effective?
Traditionally, LLM security focused on preventing single, direct attempts to elicit harmful responses. However, attackers are now employing a more nuanced strategy: multi-turn attacks. These involve iterative “probing” of the LLM, gradually building trust and subtly introducing adversarial requests over a series of interactions.
Think of it like a con artist. They don’t promptly ask for your life savings. Rather, they build rapport, establish credibility, and slowly manipulate you into lowering your guard.
Hear’s how it effectively works with LLMs:
* Establishing trust: An attacker might begin with benign queries, seemingly harmless requests for facts or creative content.
* Subtle Introduction of Harmful Requests: Once a rapport is established, the attacker subtly introduces more problematic prompts, frequently enough framed as “research,” “fictional scenarios,” or requests for roleplaying.
* Exploiting Systemic Weaknesses: LLMs are designed to detect and reject overtly malicious requests. Though, these attacks exploit systemic weaknesses that are masked in isolated interactions.By breaking down information, reassembling it, or introducing contextual ambiguity, attackers can circumvent safety guardrails.
The Cisco research demonstrates the alarming effectiveness of this approach. They found that multi-turn attacks were triumphant in eliciting disallowed content at rates two to ten times higher than single-turn attempts.
which LLMs are Most Vulnerable?
Cisco’s researchers rigorously tested a range of popular open-weight LLMs, including:
* Alibaba Qwen3-32B
* Mistral Large-2
* Meta Llama 3.3-70B-Instruct
* deepseek v3.1
* Zhipu AI GLM-4.5-Air
* Google Gemma-3-1B-1T
* Microsoft Phi-4
* OpenAI GPT-OSS-2-B
The results were concerning. Mistral Large-2 proved to be the most susceptible, with a success rate of 92.78% for multi-turn attacks. Llama 3.3 and Qwen 3 also showed high vulnerability.
Interestingly, the study revealed a correlation between model design and resilience. Google’s Gemma 3 demonstrated the most balanced performance,resisting multi-turn manipulation in 25.86% of attempts. OpenAI’s and Zhipu’s models also showed stronger resistance,rejecting over 50% of multi-turn attacks.
Why the Disparity?
The researchers attribute these differences to alignment strategies and development priorities. Models like Llama 3.3 and Qwen 3 prioritize capability – maximizing performance on a wide range of tasks. While impressive, this focus can come at the expense of robust safety guardrails.
Conversely,models like Google Gemma 3 are designed with a stronger emphasis on safety,resulting in more balanced performance.
Moreover, the study suggests that some open-weight models are released with the expectation that developers will implement their own guardrails. This places a significant burden on organizations deploying these models, potentially leaving them vulnerable if adequate security measures aren’t taken.
The Responsibility Lies with Everyone
The open-weight nature of these models – meaning anyone can download, run, and modify them – amplifies the risk. While this accessibility fosters innovation, it also means malicious actors have easy access to potentially vulnerable systems.
Addressing this challenge requires a collaborative effort from the entire AI ecosystem:
* AI Developers: Must prioritize safety alongside capability, incorporating robust guardrails into model design.
* Security Community: needs to conduct autonomous testing and develop innovative mitigation strategies.
* Organizations Deploying LLMs: Must implement layered security controls,including multi-turn testing,threat-specific mitigation,and continuous monitoring.
**Protecting Your
Worth a look
- First AI-Driven Cyberattack Triggers Response from 30 US Tech Companies
- Boeing Launches Fourth 737 MAX Assembly Line in Everett
- Rethinking Security for the Age of AI: Introducing Project Perception (archyde.com)
- Toronto Fire Chief Warns of E-Bike Dangers After Scarborough Battery Blast (archyworldys.com)