Skip to content
View nnamu-cl's full-sized avatar

Block or report nnamu-cl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
nnamu-cl/README.md

Nicholas Namusanga

GPU Kernel Engineer. CUDA kernel development, PTX/SASS-level performance tuning, 10 years of GPU compute.

I build and optimize CUDA kernels and GPU pipelines for AI and HPC workloads. I work below the framework line: kernels written by hand, SASS read in Nsight Compute, memory traffic priced in bytes per element before the first line is written, and a retro afterwards that says where the model was wrong. Ten years of low-level C++ on compute and memory bound workloads across AI compute, real-time video, simulation and numerical methods.

Highlights

cuda-orbital-sampler · Optimizing quantum compute on CUDA, Act I

A Monte Carlo sampler for hydrogen electron densities, taken from a CPU loop to 14.6 Gsamples/s on an RTX 2070 Super. The maths per sample is tiny, so the whole problem is DRAM bandwidth and latency: the random number state alone was 86% of the bytes moved in the first six kernels. Nine kernels, a prediction and falsification line before each one, a statistical verify gate that catches the bugs the visualizers on the internet ship with, and Nsight Compute reports for every step. Article in progress.

nine CUDA kernels from 0.56 to 14.6 Gsamples/s against the DRAM write roof

psiEngine · Visual quantum physics compute engine

psiEngine Vulkan rendering, GPU compute in Slang, an ECS with atom components that take (n, l, m) and draw the orbital, a typed node graph that drives component properties, and Lua scripting through sol2. Built because every kernel that produces a cloud of numbers needs a place to rotate the cloud and poke at the parameters. The CUDA sampler above is what feeds it.

Stream parallel Euclidean distance transform in CUDA: custom thresholding and normalisation kernels around NPP's PBA+ transform, run over 64 USC SIPI textures with a synthetic correctness check.

CUDA device memory management helpers: scoped allocation sessions.

Activity

activity: isometric calendar, commit habits, languages

Focus

CUDA, PTX and SASS, Nsight Compute and Systems, roofline and byte accounting, memory bound kernel design, cuRAND and Philox, Vulkan and Slang compute, low level C++.

Contact

namusanga.com · LinkedIn · @hey_namusanga

Pinned Loading

  1. cuda-npp-distance-transform cuda-npp-distance-transform Public

    Stream parallel Euclidean distance transform in CUDA, NPP PBA+ with performance tuning

    Cuda

  2. cuda-orbital-sampler cuda-orbital-sampler Public

    Monte Carlo hydrogen orbital sampler in CUDA, 9 kernels from 0.56 to 14.6 Gsamples/s, Nsight tuned

    C++

  3. psiEngine psiEngine Public

    Visual quantum physics compute engine: Vulkan rendering, GPU compute, node graph, Lua scripting

    C