Research
研究プロジェクト・論文・書籍等
- 論文
Submission from ZMM-TTS for Blizzard Challenge 2025
- #音声処理
- #音声合成
- #低資源言語
The Blizzard Challenge 2025
We address the 2025 Blizzard Challenge for Bildts, a low-resource West Frisian language variety, using ZMM-TTS, a modular multilingual TTS model with separate text-to-vec and vec-to-waveform modules. We fine-tune the model on approximately 7 hours of Bildts data and explore the three input types supported by ZMM-TTS: raw characters, IPA characters, and phoneme representations. We also compare systems built by fine-tuning two pre-trained ZMM-TTS multilingual checkpoints. To enable multi-speaker synthesis for Bildts, we hypothesize that multilingual training will be beneficial; hence, we augment training with data from a selection of ZMM-TTS pretraining languages as well as geographically related languages (Dutch, German, English, French, Portuguese, and Spanish). Due to the lack of native speakers for evaluation, we rely primarily on objective metrics to select the final system.