Abstract
Reinforcement learning (RL) is a promising approach for deriving control policies for complex systems. As we show in two control problems, the derived policies from using the Proximal Policy Optimization (PPO) and Deep Q-Network (DQN) algorithms may lack robustness guarantees. Motivated by these issues, we propose a new hybrid algorithm, which we call Hysteresis-Based RL (HyRL), augmenting an existing RL algorithm with hysteresis switching and two stages of learning. We illustrate its properties in two examples for which PPO and DQN fail.
| Original language | English |
|---|---|
| Title of host publication | 2022 American Control Conference, ACC 2022 |
| Publisher | Institute of Electrical and Electronics Engineers |
| Pages | 2663-2668 |
| Number of pages | 6 |
| ISBN (Electronic) | 9781665451963 |
| DOIs | |
| Publication status | Published - 5 Sept 2022 |
| Event | 2022 American Control Conference, ACC 2022 - Atlanta, United States Duration: 8 Jun 2022 → 10 Jun 2022 https://acc2022.a2c2.org/ |
Conference
| Conference | 2022 American Control Conference, ACC 2022 |
|---|---|
| Abbreviated title | ACC 2022 |
| Country/Territory | United States |
| City | Atlanta |
| Period | 8/06/22 → 10/06/22 |
| Internet address |
Funding
*Research by R. G. Sanfelice has been partially supported by the National Science Foundation under Grant no. ECS-1710621, Grant no. CNS-1544396, and Grant no. CNS-2039054, by the Air Force Office of Scientific Research under Grant no. FA9550-19-1-0053, Grant no. FA9550-19-1-0169, and Grant no. FA9550-20-1-0238, and by the Army Research Office under Grant no. W911NF-20-1-0253. [email protected] 2Ricardo G. Sanfelice is with the Department of Electrical and Computer Engineering, University of California, Santa Cruz, CA 95064, USA; [email protected] 3Nathan van de wouw is with the Department of Mechanical Engineering, Eindhoven University of Technology, Eindhoven, 5612 AZ, Netherlands; [email protected]
| Funders | Funder number |
|---|---|
| National Science Foundation | CNS-1544396, ECS-1710621, CNS-2039054 |
| Air Force Office of Scientific Research (AFOSR) | FA9550-19-1-0169, FA9550-19-1-0053, FA9550-20-1-0238 |
Fingerprint
Dive into the research topics of 'Hysteresis-Based RL: Robustifying Reinforcement Learning-based Control Policies via Hybrid Control'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver