
Attackers who probe large language models rarely give up after one refusal. They reframe questions, build context across multiple turns, adopt personas, and escalate their demands gradually. New research from Cisco’s AI threat intelligence team finds that the safety benchmarks used across the industry miss almost all of this behavior. The gap between published scores and observed resilience is wide enough to misrank leading models, leaving buyers and regulators with a false sense of security.
The Research Methodology
The study paired single-turn and multi-turn evaluation across 15 closed flagship models from OpenAI, Anthropic, Google, Amazon, and xAI. Testing covered roughly 30,000 single-turn prompts and nearly 7,000 multi-turn attacks spread across more than 1,400 conversations. This large-scale approach allowed researchers to compare how models perform under the kind of iterative pressure that real adversaries apply, as opposed to the isolated queries used in most safety benchmarks.
Across the entire cohort, multi-turn attack success rates climbed as high as 88% – an order of magnitude above the lowest result in the group. More importantly, single-turn and multi-turn testing produced different rankings, different failure maps, and different tail-risk profiles. A model that appears nearly invulnerable in a single-turn test can become highly exploitable once attackers are allowed to adapt their approach across a conversation.
Key Findings: Single-Turn Scores Hide Real Exposure
Every model in the cohort failed a meaningful share of multi-turn attacks. OpenAI’s GPT-5.4 jumped roughly ninefold under iterative pressure, moving from a single-turn attack success rate (ASR) in the low single digits to nearly 25%. Google’s Gemini 3 Pro climbed from about 18% to 73%. xAI’s Grok 4.1 Fast in its non-reasoning configuration topped the cohort at 88%. Anthropic’s Claude family posted the strongest single-turn refusal performance, with single-turn ASRs in the low single digits, yet still landed in the 11% to 16% range once attackers were allowed to adapt.
Cross-regime gaps ran in both directions. Gemini 3 Pro rose by more than 55 points under iterative testing. All three Amazon Nova variants moved the opposite way: Nova 2 Lite recorded a relatively high single-turn ASR and the lowest multi-turn ASR in the entire cohort at about 8%. More than half of the models tested showed an absolute gap of at least 15 points between the two regimes.
These results underscore a fundamental problem: single-turn benchmarks measure a model’s ability to resist a single, blunt attack. Real-world attackers do not operate that way. They probe for weaknesses across a conversation, learning from each refusal and refining their approach. As Cisco researchers noted, the question buyers and regulators should ask is not how a model performs on a static test, but how it holds up against multi-turn, adaptive attacks.
Strategies Used by Attackers
The research identified five strategy families that drove most of the multi-turn outcomes: role-play and persona adoption, contextual ambiguity, refusal reframing, information decomposition, and crescendo-style escalation. Within each family, the spread between the most and least exposed model was large, often approaching the full range of the chart. This means that strategy labels mostly sort which models pull apart from one another, even where average difficulty looks similar.
On the single-turn side, three procedures dominated the rankings: Imposter AI, Soft Paraphrase, and System Prompts. By content type, hate speech, profanity, and specialized advice led. Imposter AI alone outpaced the tenth-ranked procedure by a wide margin, suggesting that targeted fixes to a handful of attack surfaces could move the aggregate numbers for most models in the cohort.
These findings highlight that the attack surface for AI models is not uniform. Some strategies are far more effective than others, and models that resist one type of attack may be completely vulnerable to another. The research underscores the need for strategy-specific reporting so that buyers can understand not just a model’s average resilience, but its weaknesses in the attack patterns that matter most.
Impact of Configuration
One of the most striking findings is how much a single configuration flag can change a model’s security profile. The same Grok 4.1 Fast model with reasoning mode enabled saw its multi-turn ASR cut roughly in half – a swing of more than 40 points tied to a single capability flag. This kind of configuration-driven safety variation does not appear on any public benchmark or model card the authors reviewed. Users running the model in its default non-reasoning configuration encounter a substantially different threat profile from users who turn reasoning on.
This discovery has profound implications for enterprise deployment and regulation. If a model’s safety is heavily dependent on settings that are not documented or tested in benchmark suites, then risk assessments based on standard evaluations are incomplete. It also means that model cards and public disclosures must include information about the configuration used during testing, as well as the expected variation across common configurations.
Guardrails Reduce Risk Without Eliminating It
Production deployments typically wrap base models in additional safety layers such as content filters, output monitors, and human-in-the-loop systems. The research acknowledges that these guardrails help attenuate risk, but they do not eliminate it. The base model sets the floor on what any production system can achieve. Just as traditional software development decisions involve risk tolerance and acceptance for the code itself and all its dependencies, the same approach applies to AI development and deployment.
However, the blast radius for a rogue or misaligned AI agent has the potential to be more damaging than a typical software flaw. As AI systems become more agentic – taking actions on behalf of users – the consequences of a successful attack multiply. The research team recommends that organizations buying or deploying AI take three operational steps: publish ASR by strategy family on every model release, gate deployments on regressions in the top three procedures and content types using a 3-point threshold, and flag any model with a cross-regime gap above 15 points for manual review. Applied to this cohort, the third rule alone surfaces more than half the tested models for closer examination.
Regulatory Implications
Regulatory frameworks are beginning to point in the same direction. The NIST AI Risk Management Framework, the forthcoming NIST Cyber AI Profile (IR 8596), and Article 15 of the EU AI Act all call for adversarial robustness testing. However, none currently specify the interaction regime, strategy decomposition, or slice-support labeling that the Cisco research argues is needed for decision-grade assessment. Without such specifications, regulators may end up relying on incomplete benchmarks that fail to capture the most dangerous attack vectors.
The research adds to a growing body of evidence that multi-turn vulnerability is a structural property of current frontier models. An earlier Cisco study of eight open-weight models found that multi-turn ASR ran two to ten times higher than single-turn baselines and reached more than 90% against Mistral Large-2. This pattern holds across both open and proprietary weights, suggesting that the underlying architecture of large language models makes them inherently susceptible to iterative probing.
As AI adoption accelerates across industries, the gap between benchmark scores and real-world resilience will only become more critical. Buyers who rely on single-turn safety evaluations may be making procurement decisions based on incomplete and misleading data. The Cisco research provides a stark reminder that in the arms race between attackers and defenders, the attackers have already adapted to the multi-turn environment. It is time for the defenders to do the same.
Source:Help Net Security News
