Login (DCU Staff Only)
Login (DCU Staff Only)

DORAS | DCU Research Repository

Explore open access research and scholarly works from DCU

Advanced Search

Irish treebanking and parsing: a preliminary evaluation

Lynn, Teresa, Cetinoglu, Ozlem, Foster, Jennifer orcid logoORCID: 0000-0002-7789-4853, Uí Dhonnchadha, Elaine orcid logoORCID: 0000-0003-3448-4288, Dras, Mark orcid logoORCID: 0000-0001-9908-7182 and van Genabith, Josef orcid logoORCID: 0000-0003-1322-7944 (2012) Irish treebanking and parsing: a preliminary evaluation. In: International Conference on Linguistic Resources and Evaluation, 21-27 May 2012, Istanbul, Turkey.

Abstract
Language resources are essential for linguistic research and the development of NLP applications. Low- density languages, such as Irish, therefore lack significant research in this area. This paper describes the early stages in the development of new language resources for Irish – namely the first Irish dependency treebank and the first Irish statistical dependency parser. We present the methodology behind building our new treebank and the steps we take to leverage upon the few existing resources. We discuss language specific choices made when defining our dependency labelling scheme, and describe interesting Irish language characteristics such as prepositional attachment, copula and clefting. We manually develop a small treebank of 300 sentences based on an existing POS-tagged corpus and report an inter-annotator agreement of 0.7902. We train MaltParser to achieve preliminary parsing results for Irish and describe a bootstrapping approach for further stages of development.
Metadata
Item Type:Conference or Workshop Item (Paper)
Event Type:Conference
Refereed:Yes
Uncontrolled Keywords:Dependency; Treebank; Irish
Subjects:Computer Science > Computational linguistics
Humanities > Irish language
Humanities > Linguistics
DCU Faculties and Centres:Research Institutes and Centres > Centre for Next Generation Localisation (CNGL)
Research Institutes and Centres > National Centre for Language Technology (NCLT)
Published in: Proceedings of the International Conference on Linguistic Resources and Evaluation. .
Official URL:http://www.lrec-conf.org/proceedings/lrec2012/pdf/...
Copyright Information:© The Authors
Use License:This item is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 3.0 License. View License
Funders:Science Foundation Ireland (Grant No 07/CE/I1142) as part of the Centre for Next Generation Localisation (www.cngl.ie) at Dublin City University., German ˘ Research Foundation (Deutsche Forschungsgemeinschaft – DFG) via project D2 of SFB 732 “Incremental Specification in Context”
ID Code:17974
Deposited On:09 Apr 2013 10:30 by Jennifer Foster . Last Modified 19 Jan 2022 12:47
Documents

Full text available as:

[thumbnail of LREC2012.pdf]
Preview
PDF - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
433kB
Downloads

Downloads

Downloads per month over past year

Archive Staff Only: edit this record