Web: http://arxiv.org/abs/2209.10608

Sept. 23, 2022, 1:15 a.m. | Sara Papi, Alina Karakanta, Matteo Negri, Marco Turchi

cs.CL updates on arXiv.org arxiv.org

Speech translation for subtitling (SubST) is the task of automatically
translating speech data into well-formed subtitles by inserting subtitle breaks
compliant to specific displaying guidelines. Similar to speech translation
(ST), model training requires parallel data comprising audio inputs paired with
their textual translations. In SubST, however, the text has to be also
annotated with subtitle breaks. So far, this requirement has represented a
bottleneck for system development, as confirmed by the dearth of publicly
available SubST corpora. To fill this …

arxiv data

