all AI news
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
April 9, 2024, 4:42 a.m. | Hainan Xu, Zhehuai Chen, Fei Jia, Boris Ginsburg
cs.LG updates on arXiv.org arxiv.org
Abstract: This paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model's decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. …
abstract arxiv automatic speech recognition chinese components cs.cl cs.lg cs.sd datasets decoder eess.as embedding embeddings multiple paper pet recognition speech speech recognition text the decoder tokens type
More from arxiv.org / cs.LG updates on arXiv.org
Jobs in AI, ML, Big Data
Data Architect
@ University of Texas at Austin | Austin, TX
Data ETL Engineer
@ University of Texas at Austin | Austin, TX
Lead GNSS Data Scientist
@ Lurra Systems | Melbourne
Senior Machine Learning Engineer (MLOps)
@ Promaton | Remote, Europe
Data Scientist
@ Publicis Groupe | New York City, United States
Bigdata Cloud Developer - Spark - Assistant Manager
@ State Street | Hyderabad, India