Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages. (arXiv:2211.01338v1 [eess.AS]) | allainews.com

Nov. 3, 2022, 1:16 a.m. | Anusha Prakash, Arun Kumar, Ashish Seth, Bhagyashree Mukherjee, Ishika Gupta, Jom Kuriakose, Jordan Fernandes, K V Vikram, Mano Ranjith Kumar M, Metil

cs.CL updates on arXiv.org arxiv.org

Cross-lingual dubbing of lecture videos requires the transcription of the
original audio, correction and removal of disfluencies, domain term discovery,
text-to-text translation into the target language, chunking of text using
target language rhythm, text-to-speech synthesis followed by isochronous
lipsyncing to the original video. This task becomes challenging when the source
and target languages belong to different language families, resulting in
differences in generated audio duration. This is further compounded by the
original speaker's rhythm, especially for extempore speech. This paper …

arxiv cross-lingual dubbing pipeline scale technology videos

More from arxiv.org / cs.CL updates on arXiv.org

AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback 10 hours ago | arxiv.org

abstract application arxiv challenges +23

Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition 10 hours ago | arxiv.org

abstract applications architecture arxiv +15

Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performance 10 hours ago | arxiv.org

abstract arxiv cs.cl embeddings +15

BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models 10 hours ago | arxiv.org

abstract arxiv augmentation binary +19

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models 10 hours ago | arxiv.org

abstract arxiv bootstrap bridge +17

PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization 10 hours ago | arxiv.org

abstract arxiv conversations cs.cl +12

Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets 10 hours ago | arxiv.org

abstract arxiv asr automatic speech recognition +22

Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph 10 hours ago | arxiv.org

abstract arxiv beyond cs.ai +20

Tabular Embedding Model (TEM): Finetuning Embedding Models For Tabular RAG Applications 10 hours ago | arxiv.org

abstract applications art arxiv +26

Founding AI Engineer, Agents

@ Occam AI | New York

View on ai-jobs.net

AI Engineer Intern, Agents

@ Occam AI | US

View on ai-jobs.net

AI Research Scientist

@ Vara | Berlin, Germany and Remote

View on ai-jobs.net

Data Architect

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Data ETL Engineer

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Sr. Software Development Manager, AWS Neuron Machine Learning Distributed Training

@ Amazon.com | Cupertino, California, USA

View on ai-jobs.net