GPU Kernel Engineer. CUDA kernel development, PTX/SASS-level performance tuning, 10 years of GPU compute.
I build and optimize CUDA kernels and GPU pipelines for AI and HPC workloads. I work below the framework line: kernels written by hand, SASS read in Nsight Compute, memory traffic priced in bytes per element before the first line is written, and a retro afterwards that says where the model was wrong. Ten years of low-level C++ on compute and memory bound workloads across AI compute, real-time video, simulation and numerical methods.
cuda-orbital-sampler · Optimizing quantum compute on CUDA, Act I
A Monte Carlo sampler for hydrogen electron densities, taken from a CPU loop to 14.6 Gsamples/s on an RTX 2070 Super. The maths per sample is tiny, so the whole problem is DRAM bandwidth and latency: the random number state alone was 86% of the bytes moved in the first six kernels. Nine kernels, a prediction and falsification line before each one, a statistical verify gate that catches the bugs the visualizers on the internet ship with, and Nsight Compute reports for every step. Article in progress.
psiEngine · Visual quantum physics compute engine
Stream parallel Euclidean distance transform in CUDA: custom thresholding and normalisation kernels around NPP's PBA+ transform, run over 64 USC SIPI textures with a synthetic correctness check.
CUDA device memory management helpers: scoped allocation sessions.
CUDA, PTX and SASS, Nsight Compute and Systems, roofline and byte accounting, memory bound kernel design, cuRAND and Philox, Vulkan and Slang compute, low level C++.



