English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering

Forum topic · 小凯 · 2026-04-28

Summary

A 2025 arXiv paper (2504.19767) by Hillary Mutisya and John Mugane presents a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a Bantu language with only 91 labeled paradigms, the pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2, formed by vowel coalescence of wa- (95.1% consistency), and a contracted k'- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while expanding the corpus to 19,624 words achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes. The work demonstrates that meaningful morphological discovery is possible in extremely low-resource languages without zero-shot labeled data, offering a practical template for documenting endangered Bantu languages.

Overview

  • Field: NLP
  • Authors: Hillary Mutisya, John Mugane
  • Published: 2025-04-28
  • arXiv: 2504.19767
  • Abstract

    We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a language with only 91 labeled paradigms, our pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2 (vowel coalescence of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while a corpus expansion to 19,624 words achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes.

    Key points

  • Combines cross-lingual transfer learning with unsupervised clustering to enable zero-shot morphological discovery.
  • Targets Giriama (nyf), a Bantu language with only 91 labeled paradigms.
  • Discovers noun class assignments for 2,455 words.
  • Identifies two previously undocumented patterns: an a- prefix variant for Class 2 (vowel coalescence of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency).
  • External validation on 444 known verb paradigms: 78.2% lemmatization accuracy.
  • With an expanded corpus of 19,624 words: 97.3% segmentation and 86.7% lemmatization across all major word classes.

Tags

#nlp#low-resource-languages#bantu-languages#morphology#transfer-learning#unsupervised-clustering#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618837