A learning-based dependency to constituency conversion algorithm for the turkish language

Title A learning-based dependency to constituency conversion algorithm for the turkish language
Author Marşan, B., Yıldız, O. K., Kuzgun, A., Yenice, A, Cesur, N., Yenice, A. B., Sanıyar, E., Kuyrukçu, O., Arıcan, B. N., Yıldız, Olcay Taner
Publication Date: 2022
Publication Place - European Language Resources Association (ELRA)
Subject Constituency parsing, Constitueny dpendency conversion, Dependency parsing
Type Document
Language English
Digital Yes
Manuscript No
Library: Özyeğin University
Library Asset ID 979-109554672-6
Record ID b0f7f990-ab90-4f28-8356-600c01646dfc
Library Location Computer Science
Date 2022
Sample Text This study aims to create the very first dependency-to-constituency conversion algorithm optimised for Turkish language. For this purpose, a state-of-the-art morphologic analyser (Yıldız et al., 2019) and a feature-based machine learning model was used. In order to enhance the performance of the conversion algorithm, bootstrap aggregating meta-algorithm was integrated. While creating the conversation algorithm, typological properties of Turkish were carefully considered. A comprehensive and manually annotated UD-style dependency treebank was the input, and constituency trees were the output of the conversion algorithm. A team of linguists manually annotated a set of constituency trees. These manually annotated trees were used as the gold standard to assess the performance of the algorithm. The conversion process yielded more than 8000 constituency trees whose UD-style dependency trees are also available on GitHub. In addition to its contribution to Turkish treebank resources, this study also offers a viable and easy-to-implement conversion algorithm that can be used to generate new constituency treebanks and training data for NLP resources like constituency parsers.
View in source Özyeğin University Özyeğin University - Historical works, archives, and periodicals search engine
Özyeğin University - Historical works, archives, and periodicals search engine Özyeğin University

A learning-based dependency to constituency conversion algorithm for the turkish language

Author Marşan, B., Yıldız, O. K., Kuzgun, A., Yenice, A, Cesur, N., Yenice, A. B., Sanıyar, E., Kuyrukçu, O., Arıcan, B. N., Yıldız, Olcay Taner
Publication Date 2022
Publication Place - European Language Resources Association (ELRA)
Subject Constituency parsing, Constitueny dpendency conversion, Dependency parsing
Type Document
Language English
Digital Yes
Manuscript No
Library Özyeğin University
Library Asset ID 979-109554672-6
Record ID b0f7f990-ab90-4f28-8356-600c01646dfc
Library Location Computer Science
Date 2022
Sample Text This study aims to create the very first dependency-to-constituency conversion algorithm optimised for Turkish language. For this purpose, a state-of-the-art morphologic analyser (Yıldız et al., 2019) and a feature-based machine learning model was used. In order to enhance the performance of the conversion algorithm, bootstrap aggregating meta-algorithm was integrated. While creating the conversation algorithm, typological properties of Turkish were carefully considered. A comprehensive and manually annotated UD-style dependency treebank was the input, and constituency trees were the output of the conversion algorithm. A team of linguists manually annotated a set of constituency trees. These manually annotated trees were used as the gold standard to assess the performance of the algorithm. The conversion process yielded more than 8000 constituency trees whose UD-style dependency trees are also available on GitHub. In addition to its contribution to Turkish treebank resources, this study also offers a viable and easy-to-implement conversion algorithm that can be used to generate new constituency treebanks and training data for NLP resources like constituency parsers.
Özyeğin University - Historical works, archives, and periodicals search engine
Özyeğin University You are being redirected...

Please wait