The Problems of LLM-generated Data in Social Science Research
- ,
- Katherine Harrison,
- Irina Shklovski
- ,
- Linköping University,
- University of Copenhagen
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 145-168 (24 pages)Journal (Volume, Issue Number)
Sociologica (Volume 18, Issue 2)Publication milestones
- Published - 30/10/2024
Publication status
Published - 30/10/2024
ISSN
1971-8853Publication IDs
- Scopus: 85216275707
Abstract
Beyond being used as fast and cheap annotators for otherwise complex classification tasks, LLMs have seen a growing adoption for generating synthetic data for social science and design research. Researchers have used LLM-generated data for data augmentation and prototyping, as well as for direct analysis where LLMs acted as proxies for real human subjects. LLM-based synthetic data build on fundamentally different epistemological assumptions than previous synthetically generated data and are justified by a different set of considerations. In this essay, we explore the various ways in which LLMs have been used to generate research data and consider the underlying epistemological (and accompanying methodological) assumptions. We challenge some of the assumptions made about LLM-generated data, and we highlight the main challenges that social sciences and humanities need to address if they want to adopt LLMs as synthetic data generators.
Publication metrics
PlumX, opens in new tab
Social media
2
Captures
19
Citations
31
Mentions
2
Funding Details
This work was partially supported by the Wallenberg AI, Autonomous Systems and Soft-ware Program – Humanity and Society (WASP-HS) funded by the Marianne and Marcus Wallenberg Foundation.
FundersFunding numbers
Marianne and Marcus Wallenberg Foundation
-