Projects per year
Abstract
Colexification refers to the linguistic phenomenon where a single lexical form is used to convey multiple meanings. By studying cross-lingual colexifications, researchers have gained valuable insights into fields such as psycholinguistics and cognitive sciences (Jackson et al., 2019; Xu et al., 2020; Karjus et al., 2021; Schapper and Koptjevskaja-Tamm, 2022; François, 2022). While several multilingual colexification datasets exist, there is untapped potential in using this information to bootstrap datasets across such semantic features. In this paper, we aim to demonstrate how colexifications can be leveraged to create such crosslingual datasets. We showcase curation procedures which result in a dataset covering 142 languages across 21 language families across the world. The dataset includes ratings of concreteness and affectiveness, mapped with phonemes and phonological features. We further analyze the dataset along different dimensions to demonstrate potential of the proposed procedures in facilitating further interdisciplinary research in psychology, cognitive science, and multilingual natural language processing (NLP). Based on initial investigations, we observe that i) colexifications that are closer in concreteness/affectiveness are more likely to colexify; ii) certain initial/last phonemes are significantly correlated with concreteness/affectiveness intra language families, such as /k/ as the initial phoneme in both Turkic and Tai-Kadai correlated with concreteness, and /p/ in Dravidian and Sino-Tibetan correlated with Valence; iii) the type-to-token ratio (TTR) of phonemes are positively correlated with concreteness across several language families, while the length of phoneme segments are negatively correlated with concreteness; iv) certain phonological features are negatively correlated with concreteness across languages. The dataset is made public online for further research.
Original language | English |
---|---|
Title of host publication | ACL 2023 - 20th SIGMORPHON Workshop on Computational Morphology, Phonology, and Phonetics, CMPP 2023 |
Editors | Garrett Nicolai, Eleanor Chodroff, Cagri Coltekin, Fred Mailhot |
Number of pages | 12 |
Publisher | Association for Computational Linguistics |
Publication date | Jul 2023 |
Pages | 98-109 |
ISBN (Electronic) | 9781959429937 |
DOIs | |
Publication status | Published - Jul 2023 |
Event | 20th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology - Toronto, Canada Duration: 14 Jul 2023 → 14 Jul 2023 |
Conference
Conference | 20th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology |
---|---|
Country/Territory | Canada |
City | Toronto |
Period | 14/07/2023 → 14/07/2023 |
Fingerprint
Dive into the research topics of 'Colexifications for Bootstrapping Cross-lingual Datasets: The Case of Phonology, Concreteness, and Affectiveness'. Together they form a unique fingerprint.Projects
- 1 Active
-
Multilingual Modelling for Resource-Poor Languages
Bjerva, J., Lent, H. C., Chen, Y., Ploeger, E., Fekete, M. R. & Lavrinovics, E.
01/09/2022 → 31/08/2025
Project: Research
Prizes
-
EliteForsk- Elite Research Travel Grant 2024
Chen, Yiyi (Recipient), 26 Feb 2024
Prize: Research, education and innovation prizes