Preprint says 'BRANCH' can bypass multi-scanner AI guardrails, including some commercial products
A newly posted arXiv paper from cybersecurity company Mindgard and Lancaster University claims it found a way to bypass a type of AI safety system that many companies use as a backstop for chatbots and other large language model tools. The paper says its method, called BRANCH, can defeat so-called multi-scanner guardrails and that some of the resulting bypasses also carried over to several commercial moderation and safety products.
The paper, titled “BRANCH: Bypassing Multi-Scanner AI Guardrails,” was posted to arXiv as 2610.10742v1 and, according to the arXiv record, was submitted Oct. 7. Its authors are William Hackett and Peter Garraghan, whose affiliations are listed as Mindgard and Lancaster University. The paper is an unreviewed preprint, and the results described are the authors’ claims, not findings independently verified in this reporting.
That matters because guardrails are a standard safety and security layer around AI systems. They are used to scan prompts going into a model and responses coming out for prompt injection, jailbreaks, unsafe content and other policy violations. Major technology vendors market these tools to enterprises, and newer guardrail setups often rely on multiple scanners working together rather than a single detector, on the theory that an ensemble is harder to evade.
In the paper’s abstract, the authors wrote: “Our findings demonstrate that BRANCH achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques, while preserving semantic meaning within the bypass. We also show how bypasses generated by BRANCH transfer to 29 unseen guardrails, including 8 commercial black-box guardrails, improving attack success in some cases up to 100% with no additional optimization.”
The paper says those 120 scenarios were drawn from 20 seed prompts per target. It presents BRANCH as a branching tree search method built for multi-scanner systems, meaning systems that combine several detectors to judge whether a prompt or model output should be blocked. The authors argue that many prior attacks focused on beating one classifier at a time, while BRANCH is designed to optimize across several scanners at once.
The commercial relevance of the paper comes from its transferability claim. According to the authors, bypasses generated against one set of targets also worked against 29 of 32 unseen guardrails without additional optimization, including eight commercial black-box products. The paper explicitly names Lakera Guard, Amazon Bedrock Guardrail, Google Model Armor, Azure Content Safety, Azure Prompt Shield, Azure Foundry Guardrails, Mistral Moderation and OpenAI Omni Moderation.
The paper also includes notable caveats. Its threat model assumes an attacker can observe per-scanner confidence scores and blocking decisions while generating attacks, a level of visibility that may not exist in all production deployments. The authors also say transferability matters because some bypasses can be generated offline and then applied to stricter black-box systems. Another limitation is the selection of seed prompts: The paper says the 20 prompts per target were chosen because they were strongly blocked, which the authors describe as a source of potential selection bias.
The PDF does not include a public BRANCH code repository or an explicit responsible-disclosure timeline, according to the research report. Even with those limitations, the preprint adds to a growing body of research arguing that AI guardrails can be evaded with adversarial text changes. What makes this paper stand out, if its reported results hold up, is that it extends that claim from single scanners to multi-scanner systems that are often treated as a stronger line of defense.