How Much Does Corpus Choice Change Dependency-Distance Estimates?
Dependency-distance estimates from a single corpus are not stable language-level properties but are significantly influenced by corpus choice Cross-treebank agreement on dependency-distance estimates was moderate at best, with treebank choice accounting for ~29% of between-group variance Substituting one treebank for another reversed nearly 40% of pairwise language orderings, undermining cross-linguistic ranking reliability Dependency-length minimization (DLM) holds universally across all treeba
Analysis
TL;DR
- Dependency-distance estimates from a single corpus are not stable language-level properties but are significantly influenced by corpus choice
- Cross-treebank agreement on dependency-distance estimates was moderate at best, with treebank choice accounting for ~29% of between-group variance
- Substituting one treebank for another reversed nearly 40% of pairwise language orderings, undermining cross-linguistic ranking reliability
- Dependency-length minimization (DLM) holds universally across all treebanks (normalized ratio below 1), but ordinal rankings do not
- The qualitative DLM universal survives corpus substitution, but cross-linguistic ordinal rankings are corpus-conditioned composites of grammatical, register, and annotation factors
Why It Matters
This research challenges a foundational assumption in computational linguistics: that dependency-distance metrics are intrinsic properties of a language. For AI practitioners building cross-linguistic models or conducting typological studies, the findings imply that corpus selection can dramatically skew comparative results, potentially invalidating conclusions drawn from single-corpus analyses.
Technical Details
- Analyzed 38 same-language treebank pairs from Universal Dependencies v2.18 to compare mean dependency-distance estimates across independently compiled corpora
- Employed concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design to assess robustness across preprocessing variations
- Found that cross-treebank disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications
- Demonstrated that dependency-length minimization (normalized ratio < 1) was confirmed in every treebank, establishing DLM as a robust qualitative universal
- Quantified that treebank choice explained approximately 29% of between-group variance in dependency-distance estimates
Industry Insight
- Researchers should treat cross-linguistic dependency-distance rankings as provisional rather than definitive, and always report which corpus or corpora underpin their claims
- Multiverse or multi-corpus analysis should become standard practice when making typological generalizations, as single-corpus findings may reflect annotation conventions or register biases rather than true linguistic properties
- AI systems trained on or evaluated against dependency-distance metrics should account for corpus-specific variance, particularly in low-resource settings where only one treebank may be available
Disclaimer: The above content is generated by AI and is for reference only.