Provable Reward-Agnostic Preference-Based Reinforcement Learning | allainews.com

April 18, 2024, 4:43 a.m. | Wenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. Lee

stat.ML updates on arXiv.org arxiv.org

arXiv:2305.18505v3 Announce Type: replace-cross
Abstract: Preference-based Reinforcement Learning (PbRL) is a paradigm in which an RL agent learns to optimize a task using pair-wise preference-based feedback over trajectories, rather than explicit reward signals. While PbRL has demonstrated practical success in fine-tuning language models, existing theoretical work focuses on regret minimization and fails to capture most of the practical frameworks. In this study, we fill in such a gap between theoretical PbRL and practical algorithms by proposing a theoretical reward-agnostic PbRL …

abstract agent arxiv cs.ai cs.lg feedback fine-tuning language language models math.st paradigm practical reinforcement reinforcement learning stat.ml stat.th success type wise work

More from arxiv.org / stat.ML updates on arXiv.org

Non-asymptotic estimates for accelerated high order Langevin Monte Carlo algorithms 1 day, 19 hours ago | arxiv.org

abstract algorithms arxiv convergence +9

Entropic covariance models 2 days, 19 hours ago | arxiv.org

abstract arxiv challenges covariance +12

Bump hunting through density curvature features 2 days, 19 hours ago | arxiv.org

abstract arxiv construct data +18

Uncertainty quantification in metric spaces 2 days, 19 hours ago | arxiv.org

abstract algorithms arxiv datasets +15

Guiding adaptive shrinkage by co-data to improve regression-based prediction and feature selection 2 days, 19 hours ago | arxiv.org

abstract arxiv clinical data +17

A general error analysis for randomized low-rank approximation with application to data assimilation 2 days, 19 hours ago | arxiv.org

abstract algebra algorithms analysis +17

Calabi-Yau Four/Five/Six-folds as $\mathbb{P}^n_\textbf{w}$ Hypersurfaces: Machine Learning, Approximation, and Generation 3 days, 19 hours ago | arxiv.org

abstract approximation arxiv five +17

Bayesian Quantile Regression with Subset Selection: A Posterior Summarization Perspective 3 days, 19 hours ago | arxiv.org

abstract arxiv bayesian distribution +16

The Projected Covariance Measure for assumption-lean variable significance testing 3 days, 19 hours ago | arxiv.org

abstract arxiv covariance lean +14

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net

Research Engineer

@ Allora Labs | Remote

View on ai-jobs.net

Ecosystem Manager

@ Allora Labs | Remote

View on ai-jobs.net

Founding AI Engineer, Agents

@ Occam AI | New York

View on ai-jobs.net