Abstract
Modern engineering systems, e.g. in robotics and manufacturing, often operate under partial observability: decisions must be made from indirect, noisy, and incomplete measurements rather than full knowledge of the system state. Achieving acceptable performance using direct output-feedback policies and systematically improving these policies is challenging: performance changes may stem from noise, estimation errors, or limited data, and experimental results may not generalize beyond the observed trajectories. This thesis studies how control policies can be improved and evaluated under partial information with limited data. The central challenge is to make reliable, data-driven decisions when observations are indirect. To address this challenge, methods are developed for policy iteration, enabling principled improvement under partial observability. The first part of the thesis addresses policy iteration when full state information is unavailable and control decisions must be based solely on measured outputs. This situation arises naturally in engineering systems due to sensing limitations, noise, or high-dimensional state spaces. Classical approaches rely on belief-state representations, which are often computationally intractable. This part develops policy iteration methods that operate directly in the output space using memoryless or static output-feedback policies. The considered systems are partially observable Markov decision processes, covering both discrete and continuous spaces, including linear-quadratic systems with Gaussian noise. By carefully structuring the policy improvement and evaluation steps, it is shown that monotonic performance improvement can be achieved despite the non-Markovian nature of the output process. An optimal structure for alternating improvement and evaluation steps with the least amount of computations is identified, leading to substantial gains in computational efficiency. Furthermore, model-free variants are introduced that estimate performance measures directly from data, enabling learning without explicit system models. The resulting framework is applied to classical output-feedback control problems, demonstrating that reinforcement learning-based policy iteration provides a viable alternative to traditional analytical and gradient-based methods. This connection highlights how modern learning algorithms can be used to design output-feedback controllers for systems commonly encountered in engineering. The second part of the thesis focuses on policy evaluation for fully observable Markov decision processes, with an emphasis on data efficiency. Policy evaluation is a key component in the aforementioned policy iteration algorithms. Classical methods often require large data sets and fail to exploit prior knowledge that is frequently available in engineering applications. The thesis develops least-squares temporal-difference methods that explicitly connect data-driven value estimation to underlying probabilistic models of the system dynamics. By interpreting least-squares policy evaluation as a model-based procedure, prior information about system behaviour can be incorporated through Bayesian principles, leading to significantly improved convergence and reduced estimation error when data are limited. In addition, the role of eligibility traces in least-squares temporal-difference learning is analysed in depth. The thesis shows that, unlike in direct temporal-difference methods, extensions using expected or mixed eligibility traces do not fundamentally improve least-squares value-function estimates, and are equivalent to simpler formulations. Together, these results clarify the theoretical foundations of data-efficient policy evaluation and provide practical guidance for applying reinforcement learning methods in applications where experimental data are expensive to obtain and/or simulators or digital twins are available. The third part of the thesis focuses on state estimation for dynamical systems with noisy, nonlinear, and possibly high-dimensional outputs, addressing scenarios where decisions based solely on measured outputs are insufficient. Accurate state estimates are essential for feedback control, measurement, and monitoring tasks, but exact Bayesian methods are often impractical due to the curse of dimensionality. This part develops Bayesian estimation techniques that reduce computational complexity while retaining probabilistic rigour. In the context of vision-based pose estimation, Bayesian filtering methods are combined with Gaussian process regression and deep learning to construct reliable likelihood models from image data. These approaches are validated in both simulated and industrial environments, illustrating their effectiveness in automated metrology and manufacturing applications. In addition, a frequency-domain perspective on Bayesian filtering is introduced for linear systems with nonlinear outputs and general noise distributions. By representing probability densities using Fourier basis functions, exact Bayesian estimation can be carried out in a countable coefficient space, leading to tractable approximation schemes with controlled complexity growth. The applicability of this approach is demonstrated in the context of electron microscopy, showing that accurate state estimation can be achieved without prohibitive computational cost. Together, the three parts of this thesis develop methods for conducting policy iteration with limited data availability and partial information, addressing both the evaluation and improvement of policies as well as the estimation of hidden system states.
| Original language | English |
|---|---|
| Qualification | Doctor of Philosophy |
| Awarding Institution |
|
| Supervisors/Advisors |
|
| Award date | 2 Jun 2026 |
| Place of Publication | Eindhoven |
| Publisher | |
| Print ISBNs | 978-90-386-6700-3 |
| Publication status | Published - 2 Jun 2026 |
Bibliographical note
Proefschrift.Fingerprint
Dive into the research topics of 'Policy Iteration with Limited Data and Partial Information'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver