[R] WavJourney: Compositional Audio Creation with Large Language Models - University of Surrey 2023 | allainews.com

Aug. 25, 2023, 6:43 p.m. | /u/Singularian2501

Machine Learning www.reddit.com

Paper: [https://arxiv.org/abs/2307.14335](https://arxiv.org/abs/2307.14335)

Github: [https://github.com/Audio-AGI/WavJourney](https://github.com/Audio-AGI/WavJourney)

Project Page: [https://audio-agi.github.io/WavJourney\_demopage/](https://audio-agi.github.io/WavJourney_demopage/)

Demo: [https://huggingface.co/spaces/Audio-AGI/WavJourney](https://huggingface.co/spaces/Audio-AGI/WavJourney)

Abstract:

>Large Language Models (LLMs) have shown great promise in integrating diverse expert models to tackle intricate language and vision tasks. Despite their significance in advancing the field of Artificial Intelligence Generated Content (AIGC), their potential in intelligent audio content creation remains unexplored. In this work, we tackle the problem of creating audio content with storylines encompassing speech, music, and sound effects, guided by text instructions. We present WavJourney, a system …

abstract aigc artificial artificial intelligence audio diverse expert generated intelligence intelligent language language models large language large language models llms machinelearning music significance speech tasks vision work

More from www.reddit.com / Machine Learning

[N] AI engineers report burnout and rushed rollouts as ‘rat race’ to stay competitive hits … 9 hours ago | www.reddit.com

ai tools article artificial artificial intelligence +17

[D] software to design figures 11 hours ago | www.reddit.com

algorithms alphatensor alphazero create +11

[D] How to train a text detection model that will detect it's orientation (rotation) ranging … 11 hours ago | www.reddit.com

case convention detection image +6

[R] HGRN2: Gated Linear RNNs with State Expansion 16 hours ago | www.reddit.com

abstract attention expansion however +15

[R] A Primer on the Inner Workings of Transformer-based Language Models 16 hours ago | www.reddit.com

abstract advanced authors insights +9

[D] Fine-tune Phi-3 model for domain specific data - seeking advice and insights 19 hours ago | www.reddit.com

accuracy advice benchmark data +11

[R] Iterative Reasoning Preference Optimization 23 hours ago | www.reddit.com

iterative machinelearning optimization reasoning

[D] Good strategies / resources to improve MLOps skills as a PhD student / researcher 1 day, 4 hours ago | www.reddit.com

eventually good index industry +12

[Discussion] Should I go to ICML and present my paper? 1 day, 4 hours ago | www.reddit.com

academia data data scientist future +10

AI Engineer Intern, Agents

@ Occam AI | US

View on ai-jobs.net

AI Research Scientist

@ Vara | Berlin, Germany and Remote

View on ai-jobs.net

Data Architect

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Data ETL Engineer

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Lead GNSS Data Scientist

@ Lurra Systems | Melbourne

View on ai-jobs.net

Lead Data Modeler

@ Sherwin-Williams | Cleveland, OH, United States

View on ai-jobs.net