Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning | allainews.com

Feb. 9, 2024, 5:43 a.m. | Zhiheng Xi Wenxiang Chen Boyang Hong Senjie Jin Rui Zheng Wei He Yiwen Ding Shichun Liu

cs.LG updates on arXiv.org arxiv.org

In this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires …

benefits challenge core cs.ai cs.cl cs.lg curriculum identify language language models large language large language models novel outcome supervision paper process process supervision reasoning reinforcement reinforcement learning supervision through training

More from arxiv.org / cs.LG updates on arXiv.org

(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum 1 day, 12 hours ago | arxiv.org

abstract aim arxiv cs.lg +12

Nash Learning from Human Feedback 1 day, 12 hours ago | arxiv.org

abstract arxiv cs.ai cs.gt +20

GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs 1 day, 12 hours ago | arxiv.org

abstract arxiv become cs.cv +16

Trainwreck: A damaging adversarial attack on image classifiers 1 day, 12 hours ago | arxiv.org

adversarial arxiv classifiers cs.cr +5

Fast Controllable Diffusion Models for Undersampled MRI Reconstruction 1 day, 12 hours ago | arxiv.org

abstract acquisition arxiv cs.lg +13

MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems 1 day, 12 hours ago | arxiv.org

abstract analysis arxiv beyond +24

From Classification to Segmentation with Explainable AI: A Study on Crack Detection and Growth Monitoring 1 day, 12 hours ago | arxiv.org

abstract arxiv classification cs.cv +22

Exploring Meta Information for Audio-based Zero-shot Bird Classification 1 day, 12 hours ago | arxiv.org

abstract advances arxiv audio +22

Occlusion-Aware Deep Convolutional Neural Network via Homogeneous Tanh-transforms for Face Parsing 1 day, 12 hours ago | arxiv.org

abstract arxiv become convolutional +16

Senior Machine Learning Engineer

@ GPTZero | Toronto, Canada

View on ai-jobs.net

Software Engineer III -Full Stack Developer - ModelOps, MLOps

@ JPMorgan Chase & Co. | NY, United States

View on ai-jobs.net

Senior Lead Software Engineer - Full Stack Senior Developer - ModelOps, MLOps

@ JPMorgan Chase & Co. | NY, United States

View on ai-jobs.net

Software Engineer III - Full Stack Developer - ModelOps, MLOps

@ JPMorgan Chase & Co. | NY, United States

View on ai-jobs.net

Research Scientist (m/w/d) - Numerische Simulation Laser-Materie-Wechselwirkung

@ Fraunhofer-Gesellschaft | Freiburg, DE, 79104

View on ai-jobs.net

Research Scientist, Speech Real-Time Dialog

@ Google | Mountain View, CA, USA

View on ai-jobs.net