Jan. 10, 2022, 2:10 a.m. | Weichao Zhou, Wenchao Li

cs.LG updates on arXiv.org arxiv.org

Reward design is a fundamental problem in reinforcement learning (RL). A
misspecified or poorly designed reward can result in low sample efficiency and
undesired behaviors. In this paper, we propose the idea of programmatic reward
design, i.e. using programs to specify the reward functions in RL environments.
Programs allow human engineers to express sub-goals and complex task scenarios
in a structured and interpretable way. The challenge of programmatic reward
design, however, is that while humans can provide the high-level structures, …

arxiv design

Senior Machine Learning Engineer (MLOps)

@ Promaton | Remote, Europe

IT Commercial Data Analyst - ESO

@ National Grid | Warwick, GB, CV34 6DA

Stagiaire Data Analyst – Banque Privée - Juillet 2024

@ Rothschild & Co | Paris (Messine-29)

Operations Research Scientist I - Network Optimization Focus

@ CSX | Jacksonville, FL, United States

Machine Learning Operations Engineer

@ Intellectsoft | Baku, Baku, Azerbaijan - Remote

Data Analyst

@ Health Care Service Corporation | Richardson Texas HQ (1001 E. Lookout Drive)