Overview
Field: NLP Authors: Ramon Ferrer-i-Cancho, Catherine Hobaiter, Thore Bergman Published: 2025-07-09 arXiv: 2507.06818
Abstract (translated)
Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human languages, an unsupervised parser can be evaluated because some gold standard is usually available or can be created. For other species, a gold standard is unknown. Thus one may conclude that it is impossible to determine the accuracy of an unsupervised parser and, consequently, that dependency parsing is unfeasible in other species.
However, the authors apply recent advances in network science to demonstrate that the proportion of correct edges retrieved by a parser must be high for the sequences of vocalizations or gestures that non-human primates produce, due to the fast decay of the sequence length distribution. In contrast, human language sequences lack this property. Therefore, evaluating parsers without a gold standard is feasible in non-human primates but is a hard problem in humans.
Key Implications
- Network science results on sequence length distributions allow accuracy guarantees for unsupervised parsers on non-human primate data without any gold standard.
- This opens the door to quantitative studies of syntactic structure in animal communication systems.
- The same guarantee does not hold for human language, where gold standards remain necessary for evaluation.