Senior Data Scientist and ML Engineer, Ph.D. in Computer Engineering. I build end-to-end AI systems — from classical supervised and unsupervised models to LLM applications serving real users in production.
Right now I work on a GenAI email-processing pipeline for the Supply, Trading & Shipping division at bp, and I do postdoctoral research on machine learning and process mining at the University of Pernambuco.
What I work with
- GenAI / LLMs: RAG, LangChain, LangGraph, prompt engineering, Azure OpenAI, Amazon Bedrock, Gemini
- Machine Learning: XGBoost, LightGBM, neural networks, LSTM/GRU, time series, clustering, SHAP
- Platform: Python, SQL, AWS, Azure, Databricks, Spark, Docker, Kubernetes, MLflow
About this profile
Most of what I build is proprietary — the pipeline at bp, a RAG system in production at a Brazilian state court, a credit scoring product — so it isn't here. What is public is mostly research code from my Ph.D. and standalone examples.
A few repositories
| nano_gpt | A GPT built from scratch in PyTorch, step by step from bigram to full Transformer, with a head-to-head comparison of what each architectural piece actually buys you |
| example_playwright_lambda | Running Playwright inside AWS Lambda for serverless scraping, container and IAM setup included |
| brazilian-justice | Process mining over the Brazilian Justice event log, from a dataset I published on 4TU and Kaggle |
| discover_analytics_analysis_lawsuit | Code for a paper on predicting lawsuit duration with machine learning and process mining |
| HeuristicMiner | A from-scratch implementation of the Heuristic Miner algorithm |
Research
15+ peer-reviewed papers on machine learning and process mining — Google Scholar. Ph.D. thesis on clustering methods for the performance analysis of legal processes, with a research period at RWTH Aachen University.
Contact