解析类型学、数据集与模型架构对跨语言POS标注迁移语言排序的影响

Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging

摘要 Abstract

跨语言迁移学习是克服数据稀缺问题的宝贵工具,但选择合适的迁移语言仍是一项挑战。目前对语言类型学、训练数据以及模型架构在迁移语言选择中的具体作用尚未完全理解。我们采用整体方法,考察数据集特定特征和精细类型学特征如何影响基于形态句法特性的两种不同来源的词性标注迁移语言选择。尽管先前的研究在双语双向长短时记忆网络(bilingual biLSTMs)的背景下探讨了这些动态,但我们扩展了分析范围,将其应用于更现代的迁移学习管道:基于预训练多语言模型的零样本预测。我们训练了一系列迁移语言排名系统,并研究了不同的特征输入如何在不同架构下影响排名器性能。结果表明,词汇重叠、词种比和谱系距离是所有架构中最重要特征。我们的研究发现,结合类型学和数据集相关特征可获得最佳排名,且单独使用任一组特征也能实现良好性能。

Cross-lingual transfer learning is an invaluable tool for overcoming data scarcity, yet selecting a suitable transfer language remains a challenge. The precise roles of linguistic typology, training data, and model architecture in transfer language choice are not fully understood. We take a holistic approach, examining how both dataset-specific and fine-grained typological features influence transfer language selection for part-of-speech tagging, considering two different sources for morphosyntactic features. While previous work examines these dynamics in the context of bilingual biLSTMS, we extend our analysis to a more modern transfer learning pipeline: zero-shot prediction with pretrained multilingual models. We train a series of transfer language ranking systems and examine how different feature inputs influence ranker performance across architectures. Word overlap, type-token ratio, and genealogical distance emerge as top features across all architectures. Our findings reveal that a combination of typological and dataset-dependent features leads to the best rankings, and that good performance can be obtained with either feature group on its own.