Skip to content
View Asantewaah's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report Asantewaah

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Asantewaah/README.md

Hi, I'm Juliet πŸ‘‹

I'm a PhD researcher in Statistics at the University of Edinburgh, working at the intersection of causal inference, machine learning, and large-scale health data. My research focuses on estimating causal effects from observational healthcare and genomic data, specifically applying Collaborative Targeted Maximum Likelihood Estimation (C-TMLE) to the UK Biobank.

Before my PhD, I spent 4+ years as a data scientist and analyst, delivering data-driven insights for multinational clients including Microsoft, Unilever, P&G, and Betway. I love bridging the gap between rigorous statistical methodology and real-world impact.


πŸ”¬ What I Work On

  • Causal Inference - C-TMLE, TMLE, IPW, propensity score methods
  • Statistical Modelling - high-dimensional regression, semiparametric efficiency theory
  • Machine Learning - LASSO/elastic net, SuperLearner, ensemble methods
  • Data Visualisation - ggplot2, Tableau, Power BI, Looker
  • Health & Genomic Data - Dataloch, DecodeME, UK Biobank, observational study design, missing data

πŸ› οΈ Tech Stack

R Python SQL Julia Git Tableau


πŸ“‚ Featured Projects

A comparative simulation study benchmarking TMLE and C-TMLE estimators across low- and high-dimensional settings. Built entirely in R using glmnet, SuperLearner, and influence function-based inference. Motivated by my PhD research on causal effect estimation in genomic data.

R TMLE C-TMLE glmnet Causal Inference Simulation

As instruments get stronger, estimators that adjust for every covariate scatter while C-TMLE stays close to the oracle

Adjusting for every available covariate sounds safe, but covariates that only drive treatment make estimates unstable. In a 500-dataset simulation run on the Eddie HPC cluster, C-TMLE was 2.5Γ— more accurate than AIPW with strong instruments and nearly matched an oracle that knew the true confounders, without being told which covariates matter. Built with TMLE.jl, with an interactive Pluto notebook.

Julia TMLE.jl C-TMLE Simulation HPC

Actively extending TMLE.jl - a Julia package for Targeted Minimum Loss-Based Estimation published in the Journal of Open Source Software (2025), by integrating Collaborative TMLE (C-TMLE) estimators into the package. Successfully implemented Lasso C-TMLE with bootstrap simulation studies and test coverage. Developed in collaboration with the TARGENE research group at the University of Edinburgh.

Julia TMLE C-TMLE Causal Inference Open Source Research Software


πŸ“Š Applied Data Science

As a campaign targets high-value customers more strongly, the naive estimate drifts upwards while TMLE stays on the true effect

When a campaign targets its best customers, a naive comparison overstated the email's impact by about 60%. Checked against a real randomised experiment on 64,000 customers, AIPW and TMLE (implemented from scratch) recover the true effect with honest 95% intervals. A causal forest trained only on the targeted data then finds who responds, and its ranking holds up against the experiment.

Python TMLE AIPW Causal forest Marketing

Cumulative gains curve: calling the top-scored 30% of clients reaches 75% of subscribers

On 41,188 real bank calls, targeting the top-scored 30% of clients wins 2.5Γ— the subscriptions of random calling on the same budget, keeping 93% of the profit with 70% fewer calls. Calibrated probabilities forecast a campaign's subscriptions to within about 5% and give a break-even calling rule that needs no hindsight.

Python XGBoost SHAP Calibration Finance

Interactive R Shiny app predicting a board game's BoardGameGeek rating and what drives it

An R analysis of 15,249 BoardGameGeek games, with a live Shiny app that runs in the browser via WebAssembly. Release year is the strongest predictor, and many popular mechanics lose their advantage once year and length are held constant.

R tidyverse glmnet ranger Shiny


πŸ“š Currently

  • πŸŽ“ PhD in Statistics - University of Edinburgh (causal inference, missing data, UK Biobank)
  • πŸ”§ Integrating C-TMLE estimators into TMLE.jl (Lasso C-TMLE implemented & tested)
  • πŸ“– Reading: What If by HernΓ‘n & Robins (the causal inference bible)

πŸ“« Get In Touch

I'm always happy to connect - whether it's about causal inference, data science, or potential collaborations.

Portfolio LinkedIn Email


"Data is not just numbers, it's the story of people's lives."

Pinned Loading

  1. Asantewaah Asantewaah Public

  2. TARGENE/TMLE.jl TARGENE/TMLE.jl Public

    A Julia implementation of the Targeted Minimum Loss-based Estimation

    Julia 26 6

  3. boardgame-ratings boardgame-ratings Public

    What makes a board game highly rated? R analysis of 15,249 BoardGameGeek games, with a browser-based Shiny app that predicts a new game's rating.

    R

  4. email-campaign-causal-impact email-campaign-causal-impact Public

    Did the email campaign work? Causal inference on a targeted campaign, checked against a 64,000-customer randomised experiment. Naive vs regression, IPW, AIPW and TMLE (built from scratch), plus a c…

    Jupyter Notebook

  5. marketing-financial-analytics marketing-financial-analytics Public

    Who should a bank call? Predicting term-deposit subscriptions on 41,188 calls without leakage, then turning scores into decisions: 2.5x the subscriptions at the same budget, calibrated forecasts an…

    Jupyter Notebook

  6. ctmle-breakdown ctmle-breakdown Public

    More adjustment is not always better: when adjusting for every covariate hurts causal estimates, and how C-TMLE avoids it. A 500-dataset simulation in Julia with TMLE.jl, run on an HPC cluster, plu…

    Julia