Author
Güner, Ekrem, Mohammadi, A., Demiroğlu, Cenk
Publication Date
2012
Publication Place
-
IEEE
Subject
Speaker recognition, Speech synthesis, Statistical analysis
Type
Document
Language
English
Digital
Yes
Manuscript
No
Library
Özyeğin University
Library Asset ID
978-1-4673-1068-0
Record ID
6707d5da-edb6-44ec-8719-dd8fce587dfc
Library Location
Electrical & Electronics Engineering
Date
2012
Notes
Due to copyright restrictions, the access to the full text of this article is only available via subscription.
Sample Text
Statistical speech synthesis (SSS) approach has become one of the most popular and successful methods in the speech synthesis field. Smooth speech transitions, without the spurious errors that are observed in unit selection systems, can be generated with the SSS approach. However, a well-known issue with SSS is the lack of voice similarity to the target speaker. The issue arises both in speaker-dependent models and models that are adapted from average voices. Moreover, in speaker adaptation, similarity to the target speaker does not increase significantly after around one minute of adaptation data which potentially indicates inherent bottleneck(s) in the system. Here, we propose using the hybrid speech synthesis approach to understand the key factors behind the speaker similarity problem. To that end, we try to answer the following question: which segments and parameters of speech, if generated/synthesized better, would have a substantial improvement on speaker similarity? In this work, our hybrid methods are described and listening test results are presented and discussed.