Cultures of the AI paralinguistic in voice cloning tools.
- Ada Ada Ada,
- Stina Hasse Jørgensen,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 249-252Publication milestones
- Published - 03/06/2024
Publication status
Published - 03/06/2024
Publication IDs
- Scopus: 85198902750
Host publication title
Companion Publication of the 2024 ACM Designing Interactive Systems Conference (DIS '24 Companion)Abstract
With AI-based voice cloning tools becoming more accessible to designers, we deem it imperative to understand their paralinguistic capabilities, limitations and cultures. Paralinguistics as a feld of study is concerned with how you say something rather than what you say, and new AI-based statistical voice synthesis tools difer signifcantly from previous methods. As such, they require asking novel questions and provoking new thoughts. This paper contributes by analyzing and evaluating various voice cloning platforms by looking at how they describe their own ability to produce three diferent paralinguistic elements: laughter, stuttering and pacing. We focus on text-to-speech and hybrid approaches to voice cloning, and follow up our analyses by attempting to produce these three paralinguistic elements using the voice cloning platform ElevenLabs’ voice synthesis tools. Conclusively, we draw on our results to pose questions for further investigation into what kinds of AI paralinguistic cultures can and should be designed.
Publication metrics
PlumX, opens in new tab
Captures
4
Citations
4
Access to documents
Related Event
Title
ACM Conference on Designing Interactive Systems
Event type
ConferenceDegree of recognition
Local eventDate
01/07/2024 - 05/07/2024Location
IT University of CopenhagenCopenhagenDenmark
