AI Safety2026-07-24VentureBeat

Multi-Turn Attacks Broke AI Models 88% of Time

Cisco's AI security research team has published alarming findings: multi-turn attacks successfully broke 15 flagship AI models up to 88.3% of the time, while traditional single-turn testing methods failed to detect most vulnerabilities. In a multi-turn attack, an adversary engages the model in a back-and-forth conversation, gradually manipulating it to bypass safety guardrails, reveal sensitive information, or perform unauthorized actions. This contrasts with single-turn attacks, where a single malicious prompt is used. The research shows that current safety measures are woefully inadequate against persistent, adaptive attackers. The 15 models tested included leading large language models from various vendors, though Cisco did not publicly name all of them. The high success rate of multi-turn attacks underscores a critical blind spot in AI security: models are often evaluated in static, one-shot scenarios, but real-world interactions are dynamic and multi-step. Cisco recommends that developers implement stronger context-aware safeguards, monitor for gradual jailbreaking attempts, and adopt adversarial testing that simulates prolonged interactions. As AI models are deployed in sensitive domains like healthcare, finance, and customer service, the ability to withstand multi-turn attacks becomes not just a technical requirement but a safety imperative. This research serves as a crucial reminder that AI security must evolve alongside the capabilities of the models themselves.

Related news