Abstract
Proper parameter configuration of algorithms is essential, but often time-consuming and complex, as many parameters need to be tuned simultaneously and evaluation can be expensive. In this paper, we focus on sequential decision-making (SDM) algorithms, which are applied to problems that require a series of decisions to be taken sequentially, aiming for an optimal cumulative outcome for the agent. To do this, every time the agent needs to make a decision, SDM algorithms take the current state of the environment as input and provide a decision as output. We propose a taxonomy of algorithm configuration approaches for SDM and introduce the concept of Per-State Algorithm Configuration (PSAC). To perform PSAC automatically, we present a framework based on Reinforcement Learning (RL). We demonstrate how PSAC by RL works in practice by applying it to two SDM algorithms on two SDM problems: Monte Carlo Tree Search, to solve a collaborative order picking problem in warehouses, and AlphaZero, to play a classic board game called Connect Four. Our experiments show that, in both use cases, PSAC achieves significant performance improvements compared to fixed parameter configurations. In general, our work expands the field of automated algorithm configuration and opens new possibilities for further research on SDM algorithms and their applications. Code is available at: https://github.com/ai-for-decision-making-tue/Per-State_Algorithm_Configuration.
| Original language | English |
|---|---|
| Title of host publication | Integration of Constraint Programming, Artificial Intelligence, and Operations Research |
| Subtitle of host publication | 22nd International Conference, CPAIOR 2025, Melbourne, VIC, Australia, November 10–13, 2025, Proceedings, Part I |
| Editors | Guido Tack |
| Publisher | Springer |
| Chapter | 6 |
| Pages | 86-102 |
| Number of pages | 17 |
| ISBN (Print) | 9783031959721 |
| DOIs | |
| Publication status | Published - 2025 |
| Event | 22nd International Conference on the Integration of Constraint Programming, Artificial Intelligence, and Operations Research, CPAIOR 2025 - Melbourne, Australia Duration: 10 Nov 2025 → 13 Nov 2025 Conference number: 22 |
Publication series
| Name | Lecture Notes in Computer Science (LNCS) |
|---|---|
| Publisher | Springer |
| Volume | 15762 |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | 22nd International Conference on the Integration of Constraint Programming, Artificial Intelligence, and Operations Research, CPAIOR 2025 |
|---|---|
| Abbreviated title | CPAIOR 2025 |
| Country/Territory | Australia |
| City | Melbourne |
| Period | 10/11/25 → 13/11/25 |
Bibliographical note
Publisher Copyright:© The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
Keywords
- algorithm configuration
- reinforcement learning
- sequential decision-making
Fingerprint
Dive into the research topics of 'Algorithm Configuration in Sequential Decision-Making'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver