Deidentifying a Norwegian Clinical Corpus - an Effort to Create a Privacy-preserving Norwegian Large Clinical Language Model

Phuong Ngo; Miguel Tejedor; Therese Olsen Svenning; Taridzo Chomutare; Andrius Budrionis; Hercules Dalianis

Deidentifying a Norwegian Clinical Corpus - an Effort to Create a Privacy-preserving Norwegian Large Clinical Language Model

Phuong Ngo, Miguel Tejedor, Therese Olsen Svenning, Taridzo Chomutare, Andrius Budrionis, Hercules Dalianis

Abstract

The study discusses the methods and challenges of deidentifying and pseudonymizing Norwegian clinical text for research purposes. The results of the NorDeid tool for deidentification and pseudonymization on different types of protected health information were evaluated and discussed, as well as the extension of its functionality with regular expressions to identify specific types of sensitive information. The research used a clinical corpus of adult patients treated in a gastro-surgical department in Norway, which contains approximately nine million clinical notes. The study also highlights the challenges posed by the unique language and clinical terminology of Norway and emphasizes the importance of protecting privacy and the need for customized approaches to meet legal and research requirements.

Anthology ID:: 2024.caldpseudo-1.5
Volume:: Proceedings of the Workshop on Computational Approaches to Language Data Pseudonymization (CALD-pseudo 2024)
Month:: March
Year:: 2024
Address:: St. Julian’s, Malta
Editors:: Elena Volodina, David Alfter, Simon Dobnik, Therese Lindström Tiedemann, Ricardo Muñoz Sánchez, Maria Irena Szawerna, Xuan-Son Vu
Venues:: CALD-pseudo | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 37–43
Language:
URL:: https://aclanthology.org/2024.caldpseudo-1.5
DOI:
Bibkey:
Cite (ACL):: Phuong Ngo, Miguel Tejedor, Therese Olsen Svenning, Taridzo Chomutare, Andrius Budrionis, and Hercules Dalianis. 2024. Deidentifying a Norwegian Clinical Corpus - an Effort to Create a Privacy-preserving Norwegian Large Clinical Language Model. In Proceedings of the Workshop on Computational Approaches to Language Data Pseudonymization (CALD-pseudo 2024), pages 37–43, St. Julian’s, Malta. Association for Computational Linguistics.
Cite (Informal):: Deidentifying a Norwegian Clinical Corpus - an Effort to Create a Privacy-preserving Norwegian Large Clinical Language Model (Ngo et al., CALD-pseudo-WS 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.caldpseudo-1.5.pdf

PDF Cite Search