Skip to main navigation Skip to search Skip to main content

Sequential and Nonparametric Statistics

Course

URL study guide

https://tue.osiris-student.nl/onderwijscatalogus/extern/cursus?cursuscode=2MMS90&collegejaar=2025&taal=en

Description

The course is organized around three main topics :

1.           Survival analysis
2.           Nonparametric statistics
3.           Sequential testing

The topic survival analysis deals with settings where one observes lifetime data (if the lifetimes are in an industrial setting rather than a medical context, then we speak of reliability theory). The specific extension is that lifetimes can be censored, i.e. we cannot observe the true value of the lifetime (e.g., if the lifetime exceeds the endpoint of the data collection period). So this is another aspect of sequential statistical inference. We will see that the probability distributions of lifetimes can better be described by so-called hazard rates or failure intensities than by distribution functions or densities. We will study estimators of the distribution function and hazard rate for censored data not only when we assume a parametrized family of probability distributions, but also when we do not wish to any parametric family (so-called nonparametric estimators).
 
The second topic focuses precisely on nonparametric settings. The basic mindset of non-parametric inference is to attempt to drop stringent modeling assumptions. In contrast with parametric models, that are characterized in terms of a finite number of parameters, in nonparametric models the unknown parameters live in an infinite dimensional space (e.g., could be a class of smooth functions).  This greatly broadens the applicability of these methods as one avoids difficult-to-justify assumptions. In this course, we will focus primarily on methods for nonparametric goodness-of-fit testing with special emphasis on dealing with censored data linking to the first part of the course. We will characterize their performance of under a wide range of settings, and show that these methods are, in a certain sense, optimal, meaning there are no other methods that can significantly outperform them. The last statement is made formal through the notion of minimax optimality, and you will learn a set of techniques to characterize the fundamental difficulty of certain inference settings.

In the survival analysis topic, we encountered the situations that data becomes available over time. The topic sequential testing extends this topic by introducing statistical methods that can be used to make decisions when data becomes available sequentially. This is relevant in many practical settings, where there is a need to monitor a process and detect changes or anomalous behavior as soon as possible. More specifically, we will study the Wald’s SPRT (Sequential Probability Ratio Test), which extends the standard hypothesis testing framework to repeated testing until there are enough observations to decide between the null and alternative hypothesis. In such scenarios, there is a tradeoff not only between type I and type II errors but also the number of samples needed to reach a reliable decision. The SPRT forms a stepping stone for the statistical process monitoring framework, that intends to timely detect changes in data streams. We will study Generalized Likelihood Ratio monitoring procedures and their associated dynamic changepoint hypotheses, and study their performance in terms of run length distributions. Such procedures are widely used in manufacturing industry (usually under the outdated name SPC (Statistical Process Control), but they are also widely used in medical applications. e.g., in clinical trials to test new medicine.
 
Learning outcomes:

By the end of this course, students will be able to:
  • Identify and develop parametric and non-parametric statistical models for a wide range of settings and scenarios, and think critically about the choices and assumptions made
  • Develop sound inference procedures based on those models, that are endowed with sound performance guarantees.
  • Understand the strengths and weaknesses of the proposed models and inference procedures, and be able to use them to them in practical and real-world settings.

Objectives

“Remember that all models are wrong; the practical question is how wrong do they have to be not to be useful” is a famous quote from the statistician George Box. In the spirit of this quote, this course will introduce you to various models that go beyond the simplistic parametric random-sample and linear regression models typically encountered in introductory statistics courses. In this course you will encounter various scenarios where one must move away from very stylized models assuming hinging on normality assumptions, as well as sequential and multiple hypothesis testing scenarios and situations where one must deal with heterogeneous types of data.
 
Ultimately, the main learning goal of this course is to help you think critically about the development of useful and appropriate statistical models, and develop well founded estimation and testing procedures which are endowed with various types of performance guarantees (e.g., theoretical or asymptotic performance guarantees).

Method of Assessment

Written examination
Course period1/09/2131/08/26
Course formatCourse