Subword symmetry in natural languages
- Olga Pelloni,
- ,
- Peter Ranacher,
- Ivan Vulic,
- Tanja Samardzic
- University of Zurich,
- ,
- ,
- ,
- University of Cambridge
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 1-29 (29 pages)Journal (Volume, Issue Number)
Royal Society Open Science (Volume 12, Issue 8)Publication milestones
- Published - 21/08/2025
Publication status
Published - 21/08/2025
ISSN
2054-5703Publication IDs
- Scopus: 105013836702
Abstract
Symmetric patterns are found in the orderly arrangements of natural structures, from proteins to the symmetry in animals’ bodies. Symmetric structures are more stable and easier to describe and compress, which is why they may have been preferred as building blocks in natural selection. The idea that natural languages undergo an evolutionary process akin to the evolution of species has been pervasive in the study of language. This process might result in symmetric patterns as in other natural structures, but the notion of symmetry is rarely associated with the study of natural language. In this study, we look for symmetric patterns in text data, considering the length of subword units under a range of possible subword analyses. We study the length of subword units in 32 languages and discover that the splits of long words tend to be symmetric regardless of the segmentation method and that some automatic methods give symmetric splits at all word lengths. These results include natural language in the set of phenomena that can be described in terms of symmetry, opening a new research avenue for the empirical study of text data as a structure comparable to various other structures in the natural world.
