Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Integrating human linguistic insights into AI: theory-driven representation for multilingual text-to-speech

Cong Zhang, Huinan Zeng, Huang Liu, Jiewen Zheng

Phonetica · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Abstract This paper explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Underspecified Lexicon (FUL) as a theory-driven input representation. Unlike data-intensive end-to end models, FUL offers a compact, interpretable feature set grounded in phonological principles, enabling scalable and equitable TTS development for low-resource languages. We provide a mapping from language-specific phones to FUL feature vectors via a SAMPA intermediate and incorporate these features into a modified FastSpeech architecture. Experiments were conducted to evaluate their ability to generate native, non-native, and code-mixed speech in English and Mandarin. We ran an experiment with a small dataset and one with a larger dataset, which showed that TTS with FUL features as input could produce intelligible native speech with as little as 8 h of training data; with 100 h of training data, intelligible speech could be generated for a language not present in the training data. The approach further supports code-mixed synthesis while preserving consistent timbre and interpretable phonetic control. These results highlight the potential of theory-driven representations for building efficient, scalable, and linguistically informed TTS systems, demonstrating that phonological features can function as both analytical tools and practical inputs for speech technology.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Cong Zhang, Huinan Zeng, Huang Liu, Jiewen Zheng
Quelle
Phonetica
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
0031-8388, 1423-0321
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Cong Zhang, Huinan Zeng, Huang Liu, Jiewen Zheng (2026). Integrating human linguistic insights into AI: theory-driven representation for multilingual text-to-speech. Phonetica. https://doi.org/10.1515/phon-2025-0072
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1