B03 – Gaze dynamics and reinforcement learning during Pavlovian conditioning
Objectives
During the first two funding periods, we studied Bayesian approaches to data assimilation for dynamical cognitive models of gaze control, including model comparisons and efficient algorithms. Based on these advances, our process-based, dynamical models could reproduce and explain inter-individual differences in eye-movement control between human participants. In the third funding period, we will apply our models of eye-movement control to Pavlovian conditioning, where model parameters change over time due to reinforcement learning. While the mathematical theory of reinforcement learning contributed to our understanding of behaviours in animals and humans, our new approach will be an innovative combination of reinforcement learning with explicit, process-based models of gaze control. Thus, our work will provide an understanding of how reinforcement learning interacts with dynamical cognitive processes of gaze control such as attention allocation or working memory, and how these processes are altered in psychological disorders.
In Pavlovian conditioning, a conditioned stimulus (CS/cue, e.g., an initially neutral picture) precedes presentation of an unconditioned stimulus (US/outcome; e.g., monetary win) such that individuals learn to predict the outcome from the cue. After learning, participants respond to the cue, either focusing attention on the predictive cue (sign-trackers) or on the outcome-location (goal-trackers). We and others have demonstrated, that the long-dominant theory that dopaminergic reward prediction errors underlie Pavlovian conditioning, surprisingly is valid only for sign-trackers, while goal-trackers seem to rely on a different mechanism. Specifically, this research provided evidence that sign-trackers use model-free reinforcement learning algorithms that rely on past encounters of wins and losses via reward prediction errors and that goal-trackers use model-based reinforcement learning algorithms that involve cognitive anticipations of potential future states.
In our project, we focus on mechanisms of gaze-control in human sign-/goal-trackers. To define individuals as sign-/goal-trackers, we can assess whether a participant looked more at a visually presented cue or at the unconditioned stimulus location. Previous work on human sign-/goal-tracking has treated gaze as a summary statistic (mean looking times at relevant stimuli). However, gaze behaviour is highly dynamic and evolves sequentially, with 3 to 4 saccades (fast eye movements) occurring per second. Using process-based modelling, we will improve our understanding of control of gaze in sign- versus goal-trackers, which is essential also for clinical applications. Our resulting integrated model of gaze control and Pavlovian conditioning will predict dynamical behaviour on two time scales, (i) eye movements within trials (millisecond scale) and (ii) changes due to learning across trials (over minutes). Our mathematical models of gaze control from the previous funding periods, will be the starting point for the work in the third funding phase.
First, we will develop a computational model of individual differences in reinforcement learning across trials. Rodent data show that sign-tracker rats approach the conditioned stimulus even if this behaviour is punished via omission of the reward (a phenomenon termed negative auto-maintenance). We plan to model individual differences in such effects using Pavlovian forms of reinforcement learning based on evolutionary biases, where appetitive stimuli drive approach towards them although this behaviour is never reinforced. Second, we will perform empirical tests of key model predictions in human participants, by conducting two experimental studies on sign-/goal-tracking, involving an investigation of negative auto-maintenance and a manipulation of working memory load. Third, to explain how cues guide gaze responses, we plan to develop a novel integrated computational model for gaze dynamics during Pavlovian conditioning. We will expand our model of gaze dynamics developed in the previous funding period to capture dynamics of value learning via a reinforcement learning model. Model evaluation will be achieved by data assimilation of the novel integrated model using existing data sets and new experimental data; our approach will include hierarchical Bayesian parameter inference and model comparisons between competing models or model implementations. Resulting individual estimates of model parameters will be correlated with clinical parameters, e.g., in addiction.
Preprints
Engbert, R. and Rabe, M. M. (2023). Tutorial on dynamical modeling of eye movements in reading. doi: 10.31234/osf.io/dsvmt
Rabe, M. M., Paape, D., Mertzen, D., Vasishth, S., and Engbert, R. (2023). SEAM: An integrated activation-coupled model of sentence processing and eye movements in reading. arXiv:2303.05221
Seelig, S., Risse, S., and Engbert, R. (2020). Predictive modeling of the influence of parafoveal informationprocessing on eye guidance in reading. doi:10.31234/osf.io/vbmqn
Publications
Schwetlick, L., Reich, S., & Engbert, R. (2025). Bayesian Dynamical Modeling of Fixational Eye Movements. Biological Cybernetics, 119, 13 doi:10.1007/s00422-025-01010-8
Engbert, R., Funken, J., & Boll-Avetisyan, N. (2025). Towards Dynamical Modeling of Infants' Looking Times. WIREs Cognitive Science, 16, e70006 doi:10.1002/wcs.70006</span>
Chen, Y., Huang, D.Z., Huang, J., Reich, S., and Stuart, A.M. (2024): Efficient, Multimodal, and Derivative-Free Bayesian Inference With Fisher-Rao Gradient Flows. Inverse Problems, doi:10.1088/1361-6420/ad847b, arXiv:2406.17263
Irwin, B., and Reich, S. (2024): EnKSGD: A class of preconditioned black box optimization and inversion algorithms. SIAM Journal on Scientific Computing, 46, A2101-A2122. doi: 10.1137/23M1561142
Yadav, H., Smith, G., Reich, S., and Vasishth, S. (2023). Number feature distortion modulates cue-based retrieval in reading. Journal of Memory and Language, Vol. 129, 104400. doi: 10.1016/j.jml.2022.104400
Schwetlick, L.; Backhaus, D. & Engbert, R. (2022). A dynamical scan-path model for task-dependence during scene viewing. Psychological Review, American Psychological Association (APA). doi: 10.1037/rev0000379
Huang, D.Z., Huang, J., Reich, S., and Stuart, A.M. (2023). Efficient derivative-free Bayesian inference for large-scale inverse problems. Inverse Probelms, Vol. 38, 125006. doi: 10.1088/1361-6420/ac99fa, arXiv:2204.04386
Engbert, R., Rabe, M. M., Schwetlick, L., Seelig, S. A., Reich, S., Vasishth, S. (2022). Data assimilation in dynamical cognitive science. Trends in Cognitive Sciences, 26(2), 99-102. doi:10.1016/j.tics.2021.11.006.
Malem-Shinitski, N., Ojeda, C., and Opper, M. (2022). Variational Bayesian Inference for Nonlinear Hawkes Process with Gaussian Process Self-Effects. Entropy, 24(3), 356. doi: 10.3390/e24030356.
Rabe, M. M., Chandra, J., Krügel, A., Seelig, S. A., Vasishth, S., & Engbert, R. (2021). A Bayesian approach to dynamical modeling of eye-movement control in reading of normal, mirrored, and scrambled texts. Psychological Review doi:10.1037/rev0000268, psyarXiv
Schwetlick, L., Rothkegel, L.O.M., Trukenbrod, H.A., Engbert, R. (2020). Modeling the effects of perisaccadic attention on gaze statistics during scene viewing. Communications Biology, 3, 727. doi: 10.1038/s42003-020-01429-8
Malem-Shinitski, N., Opper, M., Reich, S., Schwetlick, S., Seelig S. A., and Engbert, R (2020): A Mathematical Model of Exploration and Exploitation in Natural Scene Viewing. PLoS Computational Biology. doi:10.1371/journal.pcbi.1007880
Seelig, S. A., Rabe, M. M., Malem-Shinitski, N., Risse, S., Reich, S., and Engbert, R. (2020). Bayesian parameter estimation for the SWIFT model of eye-movement control during reading. Journal of Mathematical Psychology, 95, 102313. doi:10.1016/j.jmp.2019.102313; arXiv: 1901.11110.
Engbert, R., Rabe, M.M., Seelig, S.A., and Reich, S. (2019): Bayesian parameter estimation for dynamical models of eye-movement control using adaptive Markov Chain Monte Carlo simulations. Forschung im HLRN-Verbund 2019.
Malem-Shinitski, N., Seelig, S. A., Reich, S. and Engbert, R. (2019): Bayesian inference for an exploration-exploitation model of human gaze control. Conference on Cognitive Computational Neuroscience, 13-16 September 2019, Berlin, Germany (extended abstract). doi:10.32470/CCN.2019.1246-0
Seelig, S. A., Rabe, M. M., Malem-Shinitski, N., Reich, S., Engbert, R. (2019). Parameter estimation for the SWIFT model of eye-movement control during reading. Conference on Cognitive Computational Neuroscience, 13-16 September 2019, Berlin, Germany (extended abstract) doi:10.32470/CCN.2019.1369-0
Schütt, H. H., Rothkegel, L. O. M., Trukenbrod, H. A., Reich, S., Wichmann, F. A. and Engbert, R. (2017): Likelihood-based parameter estimation and comparison of dynamical cognitive models. Psychological Review, 124, 505-524. doi:10.1037/rev0000068