Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
Foundational theories of persuasion from the social sciences can be used to craft adversarial prompts that circumvent LLM alignment constraints. Since LLMs are trained on large-scale human-generated text, this paper hypothesizes that they may respond more compliantly to prompts built around persuasive structures, and investigates whether models exhibit distinctive persuasive "fingerprints" in their jailbreak responses. Empirical evaluation across multiple aligned LLMs shows that persuasion-aware prompts significantly bypass safeguards, with different models showing markedly different susceptibilities to particular persuasive strategies: Vicuna and Llama2 show highly similar orderings of effective strategies, Llama3 diverges markedly, Gemma and Phi4 prioritize Authority-based framing, and DeepSeek uniquely ranks Unity as most influential. The results underscore the value of interdisciplinary, persuasion-theory-informed approaches to LLM safety research.
