Overview
- Field: NLP
- Authors: Hillary Mutisya, John Mugane
- Published: 2025-04-28
- arXiv: 2504.19767
- Combines cross-lingual transfer learning with unsupervised clustering to enable zero-shot morphological discovery.
- Targets Giriama (nyf), a Bantu language with only 91 labeled paradigms.
- Discovers noun class assignments for 2,455 words.
- Identifies two previously undocumented patterns: an a- prefix variant for Class 2 (vowel coalescence of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency).
- External validation on 444 known verb paradigms: 78.2% lemmatization accuracy.
- With an expanded corpus of 19,624 words: 97.3% segmentation and 86.7% lemmatization across all major word classes.
Abstract
We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a language with only 91 labeled paradigms, our pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2 (vowel coalescence of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while a corpus expansion to 19,624 words achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes.