B08 - Continual learning by combining reinforcement learning and data assimilation in the context of precision therapy

This project started with the second funding period in July 2021.

Understanding the basis of variability in the response of patients to drug treatment is one of
the key challenges in drug therapy, in particular for drugs with a narrow range between
ineffective and toxic doses. Precision dosing focuses on the individualisation of drug treatment
based on patient factors known to alter drug disposition and/or response. It is often combined
with measuring biomarker concentrations of a patient over time. With few exceptions, the typical
scenario is data sparse in time and with limited ability to observe the system. Model-informed
precision dosing (MIPD) leverages prior knowledge on drug pharmacokinetics and -dynamics
to support dosing decisions in the clinics. Although MIPD is typically conceptualised
within a Bayesian framework, prevailing methods tend to concentrate on dose optimisation
using maximum a-posteriori (point) estimates. Yet, this focus neglects both variability and
uncertainty, and can yield misleading estimates and thus critically affecting decision-making.
We have developed a novel class of MIPD approaches based on a combination of reinforcement
learning (RL) and sequential data assimilation (DA), and showed its applicability in the clinical
context. We aim to build on this successful development and address key points that were
identified to further substantially progress the field of MIPD, both from a theoretical as well as a
mathematical modelling and applied perspective.
In the third funding period (start Jan 2026), WP 1 aims to address lack of theoretical foundations of
MIPD for continuous learning in a hierarchical Bayesian context. To the end, we leverage developments
in gaming, autonomous driving and robotics to investigate our RL-DA approach in the language of
partially observed Markov decision processes (POMDPs), a principled framework for decision
making under uncertainty. This will allow us to transfer ideas in sampling-based approaches for
POMDPs to the specific setting of MIPD to develop efficient numerical schemes with theoretical
guarantees. WP2 addresses the important question of model choice and bias from a novel neural
network (NN) perspective. The aim is to develop NNs capable of handling truly hierarchical
data, i.e., cross-sectional samples of longitudinal data, as is the default in clinical trials. To deal
with the most important scenarios of sparse individual data, we focus on amortised in-context
approaches that allow to use prior knowledge on different levels (study, population, individual).
Importantly, the structure will be designed to flexibly adapt the dosing input, a bottleneck of
most existing approaches. WP3 addresses the medical context. In RL, the reward function is
a key ingredient that defines the optimal action (dose). In practice, rewards are often defined
based on plausible assumptions and confirmed or adapted in discussions with clinician. We want
to pursue a different route and learn from clinicians the underlying treatment strategies they
employ. The challenge is that rewards are often only indirectly observed and difficult to state
as a function of the patient’s current status. Inferring the reward from data is known as inverse
RL. Depending on the available data and whether trial-and-error strategies can be applied in
silico or in ethically acceptable experimental settings (e.g., using organoids), different inverse
RL methods are analysed with the aim to develop interpretable reward structures. A particular
challenge in medical data is its complexity and high dimensionality (e.g., in single-cell datasets)
in identifying key features and important processes. Due to our expertise, a particular focus is
on autoimmune diseases where, e.g., flare-ups are still not well understood and cannot yet be
studied using mathematical models. WP4 aims to develop novel methodologies to effectively
’mine’ real-world medical data and enhance predictive/diagnostic outcomes.