Summon a demon and bind it: A grounded theory of LLM red teaming
- ,
- ,
- Jonathan Stray
- ,
- ,
- ,
- University of Washington,
- ,
- NVIDIA
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOriginal language
EnglishArticle number
e0314658Pages from-to (Number of pages)
Pages 1-36 (36 pages)Journal (Volume, Issue Number)
PLOS ONE (Volume 20, Issue 1)Publication milestones
- Published - 15/01/2025
Publication status
Published - 15/01/2025
ISSN
1932-6203Publication IDs
- ORCID: /0000-0002-5375-9542/work/168140766
- Scopus: 85215128727
Abstract
Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a formal qualitative methodology, we interviewed dozens of practitioners from a broad range of backgrounds, all contributors to this novel work of attempting to cause LLMs to fail. We focused on the research questions of defining LLM red teaming, uncovering the motivations and goals for performing the activity, and characterizing the strategies people use when attacking LLMs. Based on the data, LLM red teaming is defined as a limit-seeking, non-malicious, manual activity, which depends highly on a team-effort and an alchemist mindset. It is highly intrinsically motivated by curiosity, fun, and to some degrees by concerns for various harms of deploying LLMs. We identify a taxonomy of 12 strategies and 35 different techniques of attacking LLMs. These findings are presented as a comprehensive grounded theory of how and why people attack large language models: LLM red teaming.
Publication metrics
PlumX, opens in new tab
Citations
10
Mentions
6
Captures
28
Funding Details
VILLUM Foundation, grant No. 37176: ATTiKA: Adaptive Tools for Technical Knowledge Acquisition. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript
FundersFunding numbers
ATTIKA
37176
