all AI news
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
April 9, 2024, 4:42 a.m. | Hainan Xu, Zhehuai Chen, Fei Jia, Boris Ginsburg
cs.LG updates on arXiv.org arxiv.org
Abstract: This paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model's decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. …
abstract arxiv automatic speech recognition chinese components cs.cl cs.lg cs.sd datasets decoder eess.as embedding embeddings multiple paper pet recognition speech speech recognition text the decoder tokens type
More from arxiv.org / cs.LG updates on arXiv.org
Jobs in AI, ML, Big Data
Data Engineer
@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania
Artificial Intelligence – Bioinformatic Expert
@ University of Texas Medical Branch | Galveston, TX
Lead Developer (AI)
@ Cere Network | San Francisco, US
Research Engineer
@ Allora Labs | Remote
Ecosystem Manager
@ Allora Labs | Remote
Founding AI Engineer, Agents
@ Occam AI | New York