system: OPERATIONAL
← back to all hacks
DEFENSE MEDIUM NEW

CASCADE: benchmarking which jailbreak defenses actually work together

A National University of Singapore study benchmarks 19 jailbreak attacks against 15 defenses and shows that layered combinations cut attack success by over 95%.

2026-09-27 // 6 min affects: llama-2, llama-3.1, vicuna, gpt-3.5, gpt-4o, gpt-5.4-mini, claude-sonnet-4, gemini-2.5-flash, gemini-3.1-flash-lite

What is this?

On 18 September 2026, researchers Jiale Luo and Eric Han, from the School of Computing at the National University of Singapore, published CASCADE Against Jailbreaks, described as the first systematic study of how jailbreak defenses perform when combined across pipeline stages rather than evaluated in isolation. Most prior jailbreak-defense research tests one technique at a time — an input filter, an output classifier, a decoding-time guard — and reports its standalone attack-success-rate (ASR) reduction under inconsistent metrics and threat models, which makes cross-paper comparison unreliable. CASCADE instead benchmarks 19 attacks against 15 defenses under one standardized, black-box, single-turn threat model, across nine target models spanning open-weight (Vicuna-7B, Llama-2-7B-chat, Llama-3.1-8B) and proprietary (GPT-3.5-turbo, GPT-4o, GPT-5.4-mini, Claude Sonnet 4, Gemini 2.5 Flash, Gemini 3.1 Flash-Lite) systems.

How it works

The attack set spans three families: white-box transferable attacks (GCG, I-GCG, AmpleGCG, DSN, Adaptive), black-box template attacks (persona jailbreaks such as DAN and AIM, disguise techniques such as sequence-breaking and code-flip encodings, output-style manipulations), and black-box LLM-assisted attacks (ReNeLLM, PAP, GPTFuzzer, TAP). The defense set covers three pipeline stages: input guards that screen a prompt before it reaches the model (perplexity filtering, WildGuard, PromptGuard), input modification techniques that rewrite or paraphrase the prompt (self-reminders, retokenization, paraphrase defenses), and output guards that screen the model’s response before it is returned (Llama Guard, output classifiers). Rather than testing each defense alone, CASCADE runs every attack against every single defense and against defense combinations layered within a stage and across stages, under a fixed query budget and a shared ASR definition so the results are directly comparable.

Standalone input guards   : 79.9% - 82.3% ASR reduction
Standalone output guards  : 67.0% - 70.8% ASR reduction
Best cross-stage stack    : 95.4% ASR reduction (-11.5% utility)
Lean utility-first stack  : 82.3% ASR reduction (-7.0% utility)

Why it matters

The finding that matters most for defenders is not a single winning defense — it is that no defense evaluated works universally, and testing defenses in isolation, the norm in most published jailbreak-defense papers, overstates what a production system will actually achieve once real traffic exercises stage interactions, false positives and latency budgets together. A team that deploys only an input classifier because it scored well in a benchmark is leaving 20-30% of residual attack success on the table that a cheap, independent output check would have caught. The 90%+ reductions CASCADE reports for layered stacks depend on which attacks are in scope: newer LLM-assisted and adaptive attacks remain harder to fully neutralize than template-based jailbreaks such as DAN, so the figures should be read as “achievable against the tested attack pool,” not as a ceiling against unknown future techniques.

Defenses

  • Layer defenses, do not pick just one. Combine an input guard with an independent output guard rather than relying on either stage alone — the paper’s cross-stage stacks consistently outperform any single-stage combination at comparable utility cost.
  • Match the stack to your risk tolerance. CASCADE’s results give concrete configurations for security-first (~95% ASR reduction, ~11% utility loss), balanced, and utility-first (~82% ASR reduction, ~7% utility loss) deployments — choose based on what a false refusal costs your product.
  • Re-test after every model or defense version bump. ASR reduction figures are tied to the specific model and defense versions evaluated in September 2026; a provider-side model update or a defense library update can shift these numbers without warning.
  • Do not assume standalone benchmark scores generalize. A defense’s published solo ASR reduction is not additive with another defense’s solo score; measure the actual combination against your own traffic and attack mix before trusting a stacked estimate.

Status

ItemStatus
PaperarXiv preprint 2609.21793, submitted 18 Sep 2026, CC BY 4.0
Peer reviewNot yet peer-reviewed (preprint stage)
Models evaluated9 total (3 open-weight, 6 proprietary), incl. GPT-5.4-mini, Claude Sonnet 4, Gemini 3.1 Flash-Lite
Attacks / defenses tested19 attacks x 15 defenses
Best reported result95.4% ASR reduction, -11.5% utility (cross-stage stack)

Figures above are drawn directly from the paper’s reported experiments; production results will vary with model version, traffic mix and attack pool.

Sources