Safety-aligned language models are designed to refuse harmful requests, but a new paper on arXiv shows that this refusal can be circumvented by placing the request inside a role-play or narrative wrapper. The same request that is refused when stated directly is answered when embedded in a story or scenario.
The paper, titled "How Narrative Wrapping Affects LLM Refusal," presents a cross-language benchmark to measure this vulnerability. The authors examine how the effect varies across languages and registers, though the available abstract does not report specific attack success rates. The work also proposes a defense, but the abstract does not describe it.
Because this is a single source, there are no conflicting findings to compare. The study highlights a practical weakness in current safety alignment: adversarial prompts need not be overtly malicious if they are framed as fiction or role-play. This suggests that safety evaluations should include narrative contexts across multiple languages.