Researchers show how explainable AI can pinpoint and break LLM safety filters

A team of computer scientists at the University of Pavia, working with a colleague at Cochin University of Science and Technology, has developed a new attack technique called XBreaking that demonstrates how the safety alignment of open-source large language models can be surgically dismantled. Rather than relying on the trial-and-error prompt crafting that characterizes most jailbreaking methods, the researchers turned the tools of explainable AI inward, using them to locate precisely where inside a transformer model the censorship mechanism lives. Their study, published in Neural Computing and Applications, answers three research questions in the affirmative: censored and uncensored models can be reliably distinguished by their internal activity, specific layers can be identified as the primary carriers of safety behavior, and injecting carefully calibrated noise into those layers can suppress the refusal mechanism while leaving the model’s general abilities largely intact.

The starting point of the work is a simple observation about how modern language models are made safe. Commercial and open-source models alike undergo a sophisticated alignment process, typically involving reinforcement learning from human feedback and supervised instruction tuning, that teaches them to refuse harmful requests ranging from malware generation to disinformation. Previous attacks on these safeguards have mostly adopted a generate-and-test strategy: attackers craft malicious prompts, observe whether the model complies, and iterate. Techniques such as the Greedy Coordinate Gradient attack, AutoDAN, PAIR and MASTERKEY all operate at the input level, appending adversarial suffixes or evolving jailbreak prompts through optimization or genetic search. These methods can be effective, but they are computationally expensive, often requiring hundreds of optimization iterations per prompt, and they treat the model as an opaque target rather than a system whose defenses can be mapped.

The researchers instead assumed a white-box threat model in which the adversary has full access to a censored model and can obtain or construct an uncensored counterpart from the same architectural family. This assumption is less restrictive than it might sound. The authors note that as of June 2026, the Hugging Face platform hosts more than five thousand repositories explicitly labeled as uncensored variants across the LLaMA, Qwen, Gemma and Mistral families, some of which have been downloaded thousands of times. Where no such variant exists, prior research has shown that safety alignment can be removed through fine-tuning with only a modest number of curated examples, allowing an adversary to synthesize a functional uncensored twin for comparison purposes.

With both a censored model and its uncensored counterpart in hand, the first stage of XBreaking is a comparative profiling exercise. For every layer of the transformer stack, the researchers computed the mean activation value and the mean attention score in response to harmful inputs, then normalized these statistics so that layers could be ranked by how much the two models diverged. The results were striking. In the censored Llama 3.2 1B model, activations peaked at the first layer and then dropped sharply in later layers, particularly the final one, while the uncensored version continued to activate normally all the way to output. Attention patterns told a similar story: the censored model showed layer-specific attention drops and spikes at deeper layers, which the authors interpret as a strategic suppression of context propagation designed to prevent harmful content from reaching the generation stage.

To move from visual inspection to a systematic method, the team framed layer identification as a binary classification problem. For each layer, they built feature vectors from the mean activation and mean attention values of both models, then used a univariate statistical technique called SelectKBest, scored with the ANOVA F-test, to find the smallest set of layers that best distinguished censored from uncensored behavior. A knee-point analysis determined the optimal number of layers for each model. The results varied revealingly by architecture. Llama models required four and eight layers respectively, consistent with alignment tuning concentrated in the final transformer blocks. Qwen models needed three and nine layers, suggesting safety filters distributed across the depth of the network. Gemma 7B and Mistral 7B required only two and one layers, which the authors attribute to aggressive weight sharing and compressed representations that concentrate alignment differences in a few high-variance components. Fingerprinting accuracy exceeded 90 percent for most models and reached 100 percent for Gemma 7B, with Mistral at 82.5 percent.

The attack itself is remarkably simple once the targets are identified. Instead of modifying the identified layers directly, XBreaking injects noise into the layer immediately preceding each optimal target, altering the input that the safety-critical component receives. Specifically, the perturbation is applied to the weight vector of the layer normalization module that follows the self-attention mechanism, which controls the scale of post-attention activations. The researchers found that in censored models, self-attention and activations are suppressed in certain layers, and that counteracting this suppression allows harmful signals to propagate. They deliberately chose the LayerNorm weight over alternatives such as the query projection, because perturbing query projections disrupts attention globally and degrades functionality rather than specifically targeting the suppression mechanism.

Calibrating the noise required careful analysis of the models’ internal weight distributions. Most LayerNorm weights lie in a narrow band between minus one and one, so the team selected noise values ranging from minus 0.75 to 0.75, large enough to shift the model’s internal state but small enough to avoid overwriting original features and destroying generation quality. Values beyond roughly one in either direction damaged the models’ ability to produce coherent text. The evaluation used the JBB-Behaviors dataset of 100 harmful behaviors across ten categories aligned with OpenAI’s usage policies, with every response manually annotated by two domain experts who achieved a Cohen’s Kappa agreement of 0.85. The optimal-balance attack success rate reached 95.35 percent for Llama 3.2 1B, 96 percent for Llama 3.1 8B, 90.24 percent for Gemma 2B and 82.76 percent for Qwen 2.5 3B, with even the smallest Qwen model reaching 77.78 percent. Mistral 7B proved most resistant at 35.56 percent, likely because its sliding-window and multi-query attention designs stabilize activations against perturbation.

Crucially, the attack preserved benign functionality. When evaluated on the HellaSwag, TruthfulQA and MMLU benchmarks, the modified models retained substantial capability, with cosine similarity to base model outputs exceeding 75 percent for several architectures and MMLU agreement rates above 68 percent for the Llama models and 92 percent for Mistral. The authors also demonstrated that the layer fingerprints generalize beyond the dataset used to discover them. When the same frozen layer selections were applied to the held-out AdvBench and HarmBench benchmarks, XBreaking achieved attack success rates with lower bounds exceeding 70 percent on Llama and Qwen models, peaking at 84.5 percent on Llama 3.1 8B, while GCG and AutoDAN reached at most 59.5 and 54 percent respectively at far greater computational cost. Most provocatively, layers identified on small models could be transferred to production-scale systems: proportionally mapped onto Llama 3.1 70B and Qwen 2.5 72B, the attack achieved success rates of up to 77.7 and 67.7 percent respectively, suggesting that safety mechanisms occupy structurally analogous positions within each model family.

The authors are explicit about the limits of their work. XBreaking is fundamentally a white-box attack that requires access to model internals, so it does not apply to commercial systems such as GPT-4, Claude or Gemini, which are exposed only as black-box APIs. All experiments were conducted offline using publicly available weights, and no external services were targeted. The team argues that full disclosure is justified because the techniques are straightforward to implement, build on established principles, and would be discoverable by any determined adversary, particularly given the widespread availability of uncensored model variants. They propose mitigations including internal activation monitoring, activation obfuscation and architectural randomization to break the deterministic layer projection that makes cross-scale transfer possible.

The broader significance of the study lies in what it reveals about the nature of safety alignment itself. Rather than being a diffuse property spread across the entire network, censorship in these models appears as a localized, fingerprintable pattern concentrated in a small number of layers, in some cases just one. That makes alignment both more legible and more fragile than many assumed. From a defensive standpoint, the same explainability tools that enable the attack could be operationalized as automated security testing, allowing developers to probe for failure modes continuously across the software lifecycle before deployment. The researchers hope their findings will push future work toward alignment techniques that are robust to weight-space modification, and toward a deeper understanding of how safety behavior is encoded in the architecture of large language models.

Subject of Research: Explainable AI analysis of security alignment vulnerabilities in large language models

Article Title: XBreaking: understanding how LLMs security alignment can be broken

Article References: Arazzi, M., Kembu, V. K., Nocera, A., & P., V. (2026). XBreaking: understanding how LLMs security alignment can be broken. Neural Computing and Applications, 38(19), Article 794. https://doi.org/10.1007/s00521-026-12511-3

Image Credits: AI Generated

DOI: 10.1007/s00521-026-12511-3

Keywords: large language models, XBreaking, explainable AI, jailbreaking, safety alignment, adversarial attacks, transformer layers, noise injection, model fingerprinting, LLM security, white-box attack, mechanistic interpretability

Tags:

 

Share Story:

Facebook
X
LinkedIn