Abstract
As applications such as augmented reality, connected vehicles, and real-time analytics grow in complexity, the need for fast and predictable task execution has become critical. Meeting these requirements is especially challenging at the network edge, where hardware is heterogeneous, power-limited, and shared among multiple co-located applications. These factors introduce performance variability, making it difficult to meet Round-Trip Time (RTT) deadlines. Modern distributed systems execute tasks across multiple servers, yet traditional schedulers and load balancers react only after performance degradation occurs, leading to inefficient resource use and possible Service Level Agreements (SLAs) violations. To address this, research has explored performance-aware decision-making, where scheduling and load balancing rely on monitoring data. However, existing methods often use coarse metrics, arbitrary subsets of system state, or high-overhead profiling that assumes homogeneous hardware. Such approaches fail to capture the rapid fluctuations and interference typical of edge clusters. Achieving strict RTT deadlines under these dynamic conditions requires anticipating performance in advance through lightweight, per-request predictions suitable for millisecond-scale workloads. This thesis develops lightweight and accurate performance predictors that enable proactive resource management in heterogeneous clusters, targeting time-sensitive applications. Instead of reacting to slowdowns, the predictors estimate task execution time -specifically RTT- while considering node heterogeneity and interference among co-located applications. These predictions offer foresight into potential bottlenecks, allowing schedulers and load balancers to place tasks on the most suitable nodes and maintain acceptable performance. The study focuses on Electron Microscopy (EM) workloads, using the Single Particle Analysis (SPA) workflow as a representative case. Monitoring data from a Kubernetes-based cluster include hundreds of Central Processing Unit (CPU), memory, and network metrics. A multi-stage methodology identifies the most relevant ones: (i) raw metrics are converted into statistical and time-series features, (ii) correlation analysis ranks their relation to RTT, and (iii) application-specific subsets are derived, revealing that each application requires a distinct set of metrics to capture variability effectively. These selected metrics are used to train lightweight machine learning models that predict both RTT and its variability at runtime. The models balance accuracy, computational cost, and delay, allowing continuous operation alongside running applications. Variability predictions guide scheduling by mapping applications to nodes that minimize interference, while RTT predictions guide load balancing by routing each request to the instance expected to respond fastest. Their performance is evaluated through simulations replicating heterogeneous edge clusters. The proposed predictors achieve high accuracy and minimal overhead through three principles: (1) selecting the most informative metrics, (2) using short observation windows to capture rapid changes, and (3) employing lightweight models for fast inference. Variability predictors reach up to 94% accuracy with prediction times under 8ms, while RTT predictors attain up to 95% accuracy with inference delays within 10% of the application RTT. Simulations show that integrating variability predictions into scheduling and RTT predictions into load balancing improves performance and reduces resource waste. These results highlight the potential of lightweight predictive models to enable runtime performance-aware orchestration in heterogeneous edge–cloud environments and form a solid basis for future deployment in production systems. The thesis is structured into chapters covering metric selection, predictive modeling, and performance-aware decision-making, and it concludes with directions for future research. This work was carried out as part of the Autonomous Distribution Architecture on Progressing Topologies and Optimization of Resources (ADAPTOR) (project number 18651), supported by the Dutch Research Council (NWO) under the Open Technology Programme.
| Original language | English |
|---|---|
| Qualification | Doctor of Philosophy |
| Awarding Institution |
|
| Supervisors/Advisors |
|
| Award date | 10 Dec 2025 |
| Place of Publication | Eindhoven |
| Publisher | |
| Print ISBNs | 978-90-386-6544-3 |
| Publication status | Published - 10 Dec 2025 |
Bibliographical note
Proefschrift.UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 12 Responsible Consumption and Production
Fingerprint
Dive into the research topics of 'Predictable Application Performance in Resource Clusters'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver