Author
Budnik, M., Besacier, L., Khodabakhsh, Ali, Demiroğlu, Cenk
Publication Date
2016
Publication Place
-
IEEE
Subject
Active learning, Annotation propagation, Clustering, Speaker identification, OCR
Type
Document
Language
English
Digital
Yes
Manuscript
No
Library
Özyeğin University
Library Asset ID
1520-6149
Record ID
4c4d72e8-2613-4318-88b4-5ad2df66ebe6
Library Location
Electrical & Electronics Engineering
Date
2016
Notes
Due to copyright restrictions, the access to the full text of this article is only available via subscription.
Sample Text
In this paper, we present an approach for minimizing human effort in manual speaker annotation. Label propagation is used at each iteration of an active learning cycle. More precisely, a selection strategy for choosing the most suitable speech track to be labeled is proposed. Four different selection strategies are evaluated and all the tracks in a corresponding cluster are gathered using agglomerative clustering in order to propagate human annotations. To further reduce the manual labor required, an optical character recognition system is used to bootstrap annotations. At each step of the cycle, annotations are used to build speaker models. The quality of the generated speaker models is evaluated at each step using an i-vector based speaker identification system. The presented approach shows promising results on the REPERE corpus with a minimum amount of human effort for annotation.
DOI
10.1109/ICASSP.2016.7472743