Speaker Characterization Using TDNN-LSTM Based Speaker Embedding

Chia Ping Chen, Su Yu Zhang, Chih Ting Yeh, Jia Ching Wang, Tenghui Wang, Chien Lin Huang

研究成果: 書貢獻/報告類型會議論文篇章同行評審

26 引文 斯高帕斯(Scopus)

摘要

In this paper we propose speaker characterization using time delay neural networks and long short-term memory neural networks (TDNN-LSTM) speaker embedding. Three types of front-end feature extraction are investigated to find good features for speaker embedding. Three kinds of data augmentation are used to increase the amount and diversity of the training data. The proposed methods are evaluated with the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) tasks. Experimental results show that the proposed methods achieve a decision cost of 0.400 with the pooled SRE 2018 development set with a single system. In addition, by applying simple average score combination on the outputs of 12 systems, the proposed methods achieve an equal error rate (EER) of 5.56% and a minimum decision cost function of 0.423 with the SRE 2016 evaluation set.

原文???core.languages.en_GB???
主出版物標題2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings
發行者Institute of Electrical and Electronics Engineers Inc.
頁面6211-6215
頁數5
ISBN(電子)9781479981311
DOIs
出版狀態已出版 - 5月 2019
事件44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Brighton, United Kingdom
持續時間: 12 5月 201917 5月 2019

出版系列

名字ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
2019-May
ISSN(列印)1520-6149

???event.eventtypes.event.conference???

???event.eventtypes.event.conference???44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019
國家/地區United Kingdom
城市Brighton
期間12/05/1917/05/19

指紋

深入研究「Speaker Characterization Using TDNN-LSTM Based Speaker Embedding」主題。共同形成了獨特的指紋。

引用此