On Evaluating Multilingual Compositional Generalization with Translated Datasets

Zi Wang; Daniel Hershcovich

doi:10.18653/v1/2023.acl-long.93

On Evaluating Multilingual Compositional Generalization with Translated Datasets

Abstract

Compositional generalization allows efficient learning and human-like inductive biases. Since most research investigating compositional generalization in NLP is done on English, important questions remain underexplored. Do the necessary compositional generalization abilities differ across languages? Can models compositionally generalize cross-lingually? As a first step to answering these questions, recent work used neural machine translation to translate datasets for evaluating compositional generalization in semantic parsing. However, we show that this entails critical semantic distortion. To address this limitation, we craft a faithful rule-based translation of the MCWQ dataset from English to Chinese and Japanese. Even with the resulting robust benchmark, which we call MCWQ-R, we show that the distribution of compositions still suffers due to linguistic divergences, and that multilingual models still struggle with cross-lingual compositional generalization. Our dataset and methodology will serve as useful resources for the study of cross-lingual compositional generalization in other tasks.

Anthology ID:: 2023.acl-long.93
Volume:: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2023
Address:: Toronto, Canada
Editors:: Anna Rogers, Jordan Boyd-Graber, Naoaki Okazaki
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1669–1687
Language:
URL:: https://aclanthology.org/2023.acl-long.93
DOI:: 10.18653/v1/2023.acl-long.93
Bibkey:
Cite (ACL):: Zi Wang and Daniel Hershcovich. 2023. On Evaluating Multilingual Compositional Generalization with Translated Datasets. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1669–1687, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):: On Evaluating Multilingual Compositional Generalization with Translated Datasets (Wang & Hershcovich, ACL 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.acl-long.93.pdf
Video:: https://aclanthology.org/2023.acl-long.93.mp4

PDF Cite Search Video