Nested Noun Phrase Identification using BERT
2024 (English)In: 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings, European Language Resources Association (ELRA) , 2024, p. 12138-12143Conference paper, Published paper (Refereed)
Abstract [en]
For several NLP tasks, an important substep is the identification of noun phrases in running text. This has typically been done by “chunking” - a way of finding minimal noun phrases by token classification. However, chunking-like methods do not represent the fact that noun phrases can be nested. This paper presents a novel method of finding all noun phrases in a sentence, nested to an arbitrary depth, using the BERT model for token classification. We show that our proposed method achieves very good results for both Swedish and English.
Place, publisher, year, edition, pages
European Language Resources Association (ELRA) , 2024. p. 12138-12143
Keywords [en]
BERT, chunking, language models, nested phrases, noun phrase, partial parsing
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:kth:diva-348778Scopus ID: 2-s2.0-85195990394OAI: oai:DiVA.org:kth-348778DiVA, id: diva2:1878688
Conference
Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024, Hybrid, Torino, Italy, May 20 2024 - May 25 2024
Note
Part of ISBN 9782493814104
QC 20240701
2024-06-272024-06-272025-02-07Bibliographically approved