all AI news
Unimodal Aggregation for CTC-based Speech Recognition
March 21, 2024, 4:48 a.m. | Ying Fang, Xiaofei Li
cs.CL updates on arXiv.org arxiv.org
Abstract: This paper works on non-autoregressive automatic speech recognition. A unimodal aggregation (UMA) is proposed to segment and integrate the feature frames that belong to the same text token, and thus to learn better feature representations for text tokens. The frame-wise features and weights are both derived from an encoder. Then, the feature frames with unimodal weights are integrated and further processed by a decoder. Connectionist temporal classification (CTC) loss is applied for training. Compared to …
abstract aggregation arxiv automatic speech recognition cs.cl cs.sd eess.as encoder feature features learn paper recognition segment speech speech recognition text token tokens type wise
More from arxiv.org / cs.CL updates on arXiv.org
Benchmarking LLMs via Uncertainty Quantification
2 days, 5 hours ago |
arxiv.org
CARE: Extracting Experimental Findings From Clinical Literature
2 days, 5 hours ago |
arxiv.org
Jobs in AI, ML, Big Data
Data Architect
@ University of Texas at Austin | Austin, TX
Data ETL Engineer
@ University of Texas at Austin | Austin, TX
Lead GNSS Data Scientist
@ Lurra Systems | Melbourne
Senior Machine Learning Engineer (MLOps)
@ Promaton | Remote, Europe
Research Scientist
@ Meta | Menlo Park, CA
Principal Data Scientist
@ Mastercard | O'Fallon, Missouri (Main Campus)