A live webcam feed with a virtual point light composited into the scene, where real objects in the camera image occlude the light based on a monocular depth estimate computed every frame on the GPU.
Move your hand in front of the camera: when it passes in front of the light the light is blocked; when it is behind, the light glows over it.
- Node.js 20+ — nodejs.org
- pnpm —
npm install -g pnpm(run once, if you don't have it) - Chrome or Edge, stable channel — this needs WebGPU, which isn't on by default in Firefox or Safari yet
- A webcam
No GPU driver setup, no CUDA, no Python. Everything — rendering and the depth model — runs through the browser's own WebGPU access to your GPU.
git clone https://github.com/AkbarSheikh-debug/depth-light-injection.git
cd depth-light-injection
pnpm install
pnpm devOpen the printed URL (usually http://localhost:5173) in Chrome or Edge
and allow camera access when prompted. The depth model (~50 MB) is fetched
from the Hugging Face CDN on first run and cached by the browser afterwards —
every run after that starts instantly.
Then press A to auto-tune the light onto whatever's in frame, and wave a
hand in front of the camera.
| Input | Effect |
|---|---|
| Mouse move | Moves the light |
Space |
Pin / unpin the light in place |
A |
Auto-tune: put the light at whatever surface is under it |
D |
Cycle debug view: composite → camera → depth → split |
S |
Save a PNG screenshot |
| Sliders | Light depth, reach, intensity, ambient, bulb size, depth influence/contrast, softness, colour, cadence |
| Dropdown (top-left) | Switch video input device |
Start by pressing A. The depth model produces relative disparity with no
metric scale, so there is no "correct" lightDepth to compute. Auto-tune reads
the actual disparity under the light and parks the light just in front of that
surface — point at yourself, press A, then move a hand in front.
To make the light do the illuminating (a dark room), pull ambient down and
push intensity up. ambient is how much of the camera image survives with
no light on it.
- Camera →
getUserMediainto a<video>, imported each frame as aGPUExternalTexture. - Depth → Depth Anything V2 Small (Apache-2.0) via Transformers.js on the
WebGPU backend, run on a 448×448 centre crop on its own cadence, decoupled
from the render loop. The result is normalized and uploaded to an
r8unormtexture. - Lighting → a single TypeGPU fullscreen fragment pass samples camera and
depth, computes radial falloff and a
smoothstepocclusion test, and composites. Recorded into oneGPUCommandEncoderper frame.
The light is treated as a real source in a pseudo-3D space, not a sprite pasted on top. Four things combine:
// 1. Pseudo-3D distance: screen offset plus a depth offset.
dist = sqrt(dot(dxy, dxy) + pow((disparity - lightDepth) * depthWeight, 2.0));
// 2. Inverse-square attenuation — an infinite tail, so light spills across the
// room instead of stopping at a hard disc edge.
atten = 1.0 / (1.0 + pow(dist / lightRadius, 2.0));
// 3. Surfaces NEARER than the light have their camera-facing side turned away
// from it, so they receive much less.
occlusion = smoothstep(lightDepth - softness, lightDepth + softness, disparity);
facing = 1.0 - 0.8 * occlusion;
// 4. Multiplicative, so the light REVEALS each surface's own colour and a dark
// room genuinely brightens. Additive just washes a flat blob over it.
rgb = cam.rgb * (ambient + lightColor * lightIntensity * atten * facing * sourceGain);sourceGain comes from bulb coverage: the depth is sampled at 37 points
spread across the bulb's own disc, and the fraction that comes back occluded
drives the dimming. Cover a quarter of the bulb and the room dims by about a
quarter; cover all of it and you keep 10%, the way cupping a hand over a lamp
leaks rather than blacking out. A single-point test would be all-or-nothing.
The current figure is shown as bulb cov in the HUD.
smoothstep rather than a hard comparison throughout, because the model's depth
boundaries are soft and a binary test aliases badly against them.
Whatever is closest to the lens claims the top of the disparity range. A desk
edge or keyboard in shot will squash everything else — including you — into the
bottom few percent, which makes lightDepth impossible to tune. Two mitigations:
the normalization window is a 2%–90% percentile (not min/max), and depth contrast applies a gamma that expands the range you actually occupy.
On an RTX 5080 in Chrome:
| Stage | Time |
|---|---|
| Render (lighting + composite) | ~0.15–0.20 ms |
| Depth inference | ~35–41 ms |
| CPU tensor convert + normalize | ~0.7–1.1 ms |
| Depth texture upload | ~0.03 ms |
| Display | 180–280 fps |
Display rate stays high because inference is decoupled — the last depth texture is reused until a new one arrives. But inference itself is ~5× slower than the 8 ms target. Transformers.js runs the ONNX graph through onnxruntime-web, which reports that some nodes fall back to CPU. There is also an unavoidable GPU→CPU→GPU round trip, since the tensor comes back on the CPU.
Closing that gap is Phase 6: reimplementing the model as hand-written TypeGPU compute kernels (~250 dispatches) so the depth buffer never leaves the GPU and joins the same command encoder as the lighting pass. That is a separate, multi-week project — build a per-layer numerical diff harness against ONNXRuntime first.
| File | Role |
|---|---|
src/main.ts |
Orchestration, GUI, stats, input |
src/renderer.ts |
TypeGPU root, shaders, bind groups, command encoder |
src/depth.ts |
Model loading, inference, normalization |
src/camera.ts |
getUserMedia + device enumeration |
src/params.ts |
Shared uniform schema and presets |
NOTES.md |
Debugging log — read this before deep-diving a bug |
- Depth Anything V2 Small only. Base and Large are CC-BY-NC.
- A
GPUExternalTextureexpires at the end of the frame's task: the import and the bind group referencing it are both rebuilt every frame. Never cache them.
Built by @AkbarSheikh-debug.
MIT — use it, fork it, ship it. The one condition is that the
copyright notice in LICENSE stays in copies and substantial
reuses of this code; that's not a suggestion, it's the actual license term.
If you build on this, a link back or a mention is appreciated but not
required beyond that.
