Synthesizing speech with selected perceptual voice qualities
- A case study with creaky voice
Frederik Rautenberg1, Fritz Seebauer2, Jana Wiechmann2, Michael Kuhlmann1,
Petra Wagner2, Reinhold Haeb-Umbach1
1Department of Communications Engineering, Paderborn University
Email: rautenberg@nt.upb.de, kuhlmann@nt.upb.de, haeb@nt.upb.de
2Phonetics Work Group, Bielefeld University
Email: fritz.seebauer@uni-bielefeld.de, jana.wiechmann@uni-bielefeld.de, petra.wagner@uni-bielefeld.de
Abstract: The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipulated and manipulated speech probes can serve to illustrate phonetic concepts that are otherwise difficult to grasp. Here, we show that a TTS system, that is augmented with a global speaker attribute manipulation block based on normalizing flows, is capable of correctly manipulating the non-persistent, localized quality of creaky voice, thus avoiding the necessity of a, typically unreliable, frame-wise creak predictor. Subjective listening tests confirm successful creak manipulation at a slightly reduced MOS score compared to the original recording.
TTS-System
Example 1
Original voice

ã = -1.5
ã = -1.0
ã = 0.0
ã = 1.0
ã = 1.5
Example 2
Original voice

Creak condition
ã = -1.5
ã = -1.0
ã = 0.0
ã = 1.0
ã = 1.5
Example 3
Original voice

Creak condition
ã = -1.5
ã = -1.0
ã = 0.0
ã = 1.0
ã = 1.5
Example 4
Original voice

Creak condition
ã = -1.5
ã = -1.0
ã = 0.0
ã = 1.0
ã = 1.5