I'm a PhD researcher in Statistics at the University of Edinburgh, working at the intersection of causal inference, machine learning, and large-scale health data. My research focuses on estimating causal effects from observational healthcare and genomic data, specifically applying Collaborative Targeted Maximum Likelihood Estimation (C-TMLE) to the UK Biobank.
Before my PhD, I spent 4+ years as a data scientist and analyst, delivering data-driven insights for multinational clients including Microsoft, Unilever, P&G, and Betway. I love bridging the gap between rigorous statistical methodology and real-world impact.
- Causal Inference - C-TMLE, TMLE, IPW, propensity score methods
- Statistical Modelling - high-dimensional regression, semiparametric efficiency theory
- Machine Learning - LASSO/elastic net, SuperLearner, ensemble methods
- Data Visualisation - ggplot2, Tableau, Power BI, Looker
- Health & Genomic Data - Dataloch, DecodeME, UK Biobank, observational study design, missing data
A comparative simulation study benchmarking TMLE and C-TMLE estimators across low- and high-dimensional settings. Built entirely in R using glmnet, SuperLearner, and influence function-based inference. Motivated by my PhD research on causal effect estimation in genomic data.
R TMLE C-TMLE glmnet Causal Inference Simulation
Adjusting for every available covariate sounds safe, but covariates that only drive treatment make estimates unstable. In a 500-dataset simulation run on the Eddie HPC cluster, C-TMLE was 2.5Γ more accurate than AIPW with strong instruments and nearly matched an oracle that knew the true confounders, without being told which covariates matter. Built with TMLE.jl, with an interactive Pluto notebook.
Julia TMLE.jl C-TMLE Simulation HPC
Actively extending TMLE.jl - a Julia package for Targeted Minimum Loss-Based Estimation published in the Journal of Open Source Software (2025), by integrating Collaborative TMLE (C-TMLE) estimators into the package. Successfully implemented Lasso C-TMLE with bootstrap simulation studies and test coverage. Developed in collaboration with the TARGENE research group at the University of Edinburgh.
Julia TMLE C-TMLE Causal Inference Open Source Research Software
When a campaign targets its best customers, a naive comparison overstated the email's impact by about 60%. Checked against a real randomised experiment on 64,000 customers, AIPW and TMLE (implemented from scratch) recover the true effect with honest 95% intervals. A causal forest trained only on the targeted data then finds who responds, and its ranking holds up against the experiment.
Python TMLE AIPW Causal forest Marketing
On 41,188 real bank calls, targeting the top-scored 30% of clients wins 2.5Γ the subscriptions of random calling on the same budget, keeping 93% of the profit with 70% fewer calls. Calibrated probabilities forecast a campaign's subscriptions to within about 5% and give a break-even calling rule that needs no hindsight.
Python XGBoost SHAP Calibration Finance
An R analysis of 15,249 BoardGameGeek games, with a live Shiny app that runs in the browser via WebAssembly. Release year is the strongest predictor, and many popular mechanics lose their advantage once year and length are held constant.
R tidyverse glmnet ranger Shiny
- π PhD in Statistics - University of Edinburgh (causal inference, missing data, UK Biobank)
- π§ Integrating C-TMLE estimators into
TMLE.jl(Lasso C-TMLE implemented & tested) - π Reading: What If by HernΓ‘n & Robins (the causal inference bible)
I'm always happy to connect - whether it's about causal inference, data science, or potential collaborations.
"Data is not just numbers, it's the story of people's lives."



