Web: http://arxiv.org/abs/2209.10887

Sept. 23, 2022, 1:15 a.m. | Haohan Guo, Fenglong Xie, Frank K. Soong, Xixin Wu, Helen Meng

cs.CL updates on arXiv.org arxiv.org

We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance
neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE)
based feature analyzer is used to encode Mel spectrograms of speech training
data by down-sampling progressively in multiple stages into MSMC
Representations (MSMCRs) with different time resolutions, and quantizing them
with multiple VQ codebooks, respectively. Multi-stage predictors are trained to
map the input text sequence to MSMCRs progressively by minimizing a combined
loss of the reconstruction Mean Square Error (MSE) and "triplet loss". …

arxiv performance stage tts

More from arxiv.org / cs.CL updates on arXiv.org

Research Scientists

@ ODU Research Foundation | Norfolk, Virginia

Embedded Systems Engineer (Robotics)

@ Neo Cybernetica | Bedford, New Hampshire

2023 Luis J. Alvarez and Admiral Grace M. Hopper Postdoc Fellowship in Computing Sciences

@ Lawrence Berkeley National Lab | San Francisco, CA

Senior Manager Data Scientist

@ NAV | Remote, US

Senior AI Research Scientist

@ Earth Species Project | Remote anywhere

Research Fellow- Center for Security and Emerging Technology (Multiple Opportunities)

@ University of California Davis | Washington, DC

Staff Fellow - Data Scientist

@ U.S. FDA/Center for Devices and Radiological Health | Silver Spring, Maryland

Staff Fellow - Senior Data Engineer

@ U.S. FDA/Center for Devices and Radiological Health | Silver Spring, Maryland

Machine Learning Data Engineer Intern (Jyoti Dharna)

@ Benson Hill | St. Louis, Missouri

Software Engineer / SDE I, Chime SDK Video Research Engineering

@ Amazon.com | East Palo Alto, California, USA

IND (New) Senior ML Ops Engineer - WiQ

@ Quantium | Hyderabad

Data Engineer

@ LendingTree | Remote