Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation | allainews.com

March 11, 2024, 4:41 a.m. | Xiaoying Zhang, Jean-Francois Ton, Wei Shen, Hongning Wang, Yang Liu

cs.LG updates on arXiv.org arxiv.org

arXiv:2403.05171v1 Announce Type: new
Abstract: We introduce Adversarial Policy Optimization (AdvPO), a novel solution to the pervasive issue of reward over-optimization in Reinforcement Learning from Human Feedback (RLHF) for Large Language Models (LLMs). Over-optimization occurs when a reward model serves as an imperfect proxy for human preference, and RL-driven policy optimization erroneously exploits reward inaccuracies. In this paper, we begin by introducing a lightweight way to quantify uncertainties in rewards, relying solely on the last layer embeddings of the reward …

abstract adversarial arxiv cs.ai cs.lg feedback human human feedback issue language language models large language large language models llms novel optimization policy reinforcement reinforcement learning reward model rlhf solution type uncertainty via

More from arxiv.org / cs.LG updates on arXiv.org

(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum 1 day, 2 hours ago | arxiv.org

abstract aim arxiv cs.lg +12

Nash Learning from Human Feedback 1 day, 2 hours ago | arxiv.org

abstract arxiv cs.ai cs.gt +20

GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs 1 day, 2 hours ago | arxiv.org

abstract arxiv become cs.cv +16

Trainwreck: A damaging adversarial attack on image classifiers 1 day, 2 hours ago | arxiv.org

adversarial arxiv classifiers cs.cr +5

Fast Controllable Diffusion Models for Undersampled MRI Reconstruction 1 day, 2 hours ago | arxiv.org

abstract acquisition arxiv cs.lg +13

MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems 1 day, 2 hours ago | arxiv.org

abstract analysis arxiv beyond +24

From Classification to Segmentation with Explainable AI: A Study on Crack Detection and Growth Monitoring 1 day, 2 hours ago | arxiv.org

abstract arxiv classification cs.cv +22

Exploring Meta Information for Audio-based Zero-shot Bird Classification 1 day, 2 hours ago | arxiv.org

abstract advances arxiv audio +22

Occlusion-Aware Deep Convolutional Neural Network via Homogeneous Tanh-transforms for Face Parsing 1 day, 2 hours ago | arxiv.org

abstract arxiv become convolutional +16

Senior Machine Learning Engineer

@ GPTZero | Toronto, Canada

View on ai-jobs.net

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

View on ai-jobs.net

Sr Business Intelligence Analyst

@ T. Rowe Price | Baltimore, MD

View on ai-jobs.net

Business Intelligence Analyst, Market Insights and Analytics

@ Morningstar | Mumbai

View on ai-jobs.net

Senior Back-End Developer - Generative AI

@ Aptiv | POL Krakow - Eng

View on ai-jobs.net

System Architect (Document AI)

@ Trafigura | London - Traf Office

View on ai-jobs.net