Finite sample properties of sequential designs
Adaptive designs can make clinical trials more flexible by utilising results accumulating in the trial to modify the trial’s course in accordance with pre-specified rules. One of the most common adaptive designs are (group-) sequential designs that allow stopping a trial early, either because of overwhelming benefit or because of insufficiently promising results (futility).
These decisions are made by comparing the test statistics comparing an active treatment against a common control at each stage to futility and efficacy bounds defined under the assumption that these test statistics are multivariate Gaussian and that the variances in the treatment and control arms are ‘similar’. This means that the data are often assumed to be Gaussian with known variance and/or that the sample size is large enough for the central limit theorem to take effect. The latter typically holds true in phase III trials, as large sample sizes are commonly used in that context. In other settings, such as pre-clinical animal studies or early clinical studies, smaller sample sizes are the norm. In this project focus will be on binary outcomes.
The aim of this project is threefold:
- To evaluate the impact of this assumption in small samples in simulations
- To develop a finite sample correction for binary outcome
- To compare the operating characteristics of the correction by means of
Monte Carlo simulations.
Evaluating the Robustness and Assumptions of the Permutability Test for Predicting Individual Treatment Effects (PITE)
Artificial intelligence decision support systems in personalized medicine require rigorous validation prior to clinical integration. Although predicting patient-specific treatment responses can optimize therapeutic strategies, high-stakes medical decisions demand strict statistical reliability--particularly when evaluating patient profiles underrepresented in training distributions. Establishing robust statistical metrics to detect true treatment effect heterogeneity under constrained or out-of-distribution training data is a critical prerequisite for implementing safe, evidence-based clinical algorithms.
This project evaluates the methodological assumptions, performance bounds, and failure modes of the permutability test--a non-parametric statistical method used to detect treatment effect heterogeneity. Predicting Individual Treatment Effects (PITE) estimates counterfactual treatment effects at the patient level, advancing beyond Average Treatment Effects (ATE). However, valid PITE estimation requires objective prior evidence that treatment effect heterogeneity exists. The student will systematically evaluate the robustness of the permutability test across complex, noisy, and high-dimensional data environments to establish the boundary conditions required for valid PITE inference.
The aim of the project is to evaluate operational limits of the permutability test and the associated predictive models under varying data conditions, in particular:
- Model Complexity vs. Sample Size: Compare predictive performance across different algorithmic frameworks (standard linear models, Lasso, Elastic Net, and Deep Learning) to determine if highly parameterized models require specific sample size thresholds to satisfy baseline statistical assumptions.
- Dimensionality and Multiple Treatments: Analyze the performance of the T-learning procedure and the permutability test in high-dimensional spaces and in scenarios involving multiple treatment arms.
- Imputation Impact: Quantify how imputation error propagates through the permutability test and affects the detection of treatment effect heterogeneity.