Safe Exploration Using Bayesian World Models and Log-Barrier Optimization | allainews.com

May 10, 2024, 4:41 a.m. | Yarden As, Bhavya Sukhija, Andreas Krause

cs.LG updates on arXiv.org arxiv.org

arXiv:2405.05890v1 Announce Type: new
Abstract: A major challenge in deploying reinforcement learning in online tasks is ensuring that safety is maintained throughout the learning process. In this work, we propose CERL, a new method for solving constrained Markov decision processes while keeping the policy safe during learning. Our method leverages Bayesian world models and suggests policies that are pessimistic w.r.t. the model's epistemic uncertainty. This makes CERL robust towards model inaccuracies and leads to safe exploration during learning. In our …

abstract arxiv bayesian challenge cs.ai cs.lg decision exploration major markov optimization policy process processes reinforcement reinforcement learning safe safety tasks type while work world world models

More from arxiv.org / cs.LG updates on arXiv.org

Bypassing the Safety Training of Open-Source LLMs with Priming Attacks 17 hours ago | arxiv.org

arxiv attacks cs.ai cs.cl +7

Variational Mode Decomposition-Based Nonstationary Coherent Structure Analysis for Spatiotemporal Data 17 hours ago | arxiv.org

abstract analysis and analysis arxiv +12

Differentially private projection-depth-based medians 17 hours ago | arxiv.org

abstract arxiv cost cs.cr +19

Unified Binary and Multiclass Margin-Based Classification 17 hours ago | arxiv.org

abstract algorithms analysis and analysis +15

An Experimental Design for Anytime-Valid Causal Inference on Multi-Armed Bandits 17 hours ago | arxiv.org

abstract arxiv causal causal inference +12

Convergence of flow-based generative models via proximal gradient descent in Wasserstein space 17 hours ago | arxiv.org

abstract advantages analysis arxiv +23

Identifying the Risks of LM Agents with an LM-Emulated Sandbox 17 hours ago | arxiv.org

abstract advances agents amplify +22

Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs 17 hours ago | arxiv.org

arxiv cs.ai cs.cl cs.lg +6

Robust Online Learning over Networks 17 hours ago | arxiv.org

abstract agent agents arxiv +25

Software Engineer for AI Training Data (School Specific)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Python)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Tier 2)

@ G2i Inc | Remote

View on ai-jobs.net

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net