Skip to content

Repository files navigation

Depth-Aware Light Injection

A live webcam feed with a virtual point light composited into the scene, where real objects in the camera image occlude the light based on a monocular depth estimate computed every frame on the GPU.

Move your hand in front of the camera: when it passes in front of the light the light is blocked; when it is behind, the light glows over it.

Demo: a hand passing in front of the virtual light, partially occluding it

Requirements

  • Node.js 20+ — nodejs.org
  • pnpm — npm install -g pnpm (run once, if you don't have it)
  • Chrome or Edge, stable channel — this needs WebGPU, which isn't on by default in Firefox or Safari yet
  • A webcam

No GPU driver setup, no CUDA, no Python. Everything — rendering and the depth model — runs through the browser's own WebGPU access to your GPU.

Quick start

git clone https://github.com/AkbarSheikh-debug/depth-light-injection.git
cd depth-light-injection
pnpm install
pnpm dev

Open the printed URL (usually http://localhost:5173) in Chrome or Edge and allow camera access when prompted. The depth model (~50 MB) is fetched from the Hugging Face CDN on first run and cached by the browser afterwards — every run after that starts instantly.

Then press A to auto-tune the light onto whatever's in frame, and wave a hand in front of the camera.

Controls

Input Effect
Mouse move Moves the light
Space Pin / unpin the light in place
A Auto-tune: put the light at whatever surface is under it
D Cycle debug view: composite → camera → depth → split
S Save a PNG screenshot
Sliders Light depth, reach, intensity, ambient, bulb size, depth influence/contrast, softness, colour, cadence
Dropdown (top-left) Switch video input device

Start by pressing A. The depth model produces relative disparity with no metric scale, so there is no "correct" lightDepth to compute. Auto-tune reads the actual disparity under the light and parks the light just in front of that surface — point at yourself, press A, then move a hand in front.

To make the light do the illuminating (a dark room), pull ambient down and push intensity up. ambient is how much of the camera image survives with no light on it.

How it works

  1. Camera → getUserMedia into a <video>, imported each frame as a GPUExternalTexture.
  2. Depth → Depth Anything V2 Small (Apache-2.0) via Transformers.js on the WebGPU backend, run on a 448×448 centre crop on its own cadence, decoupled from the render loop. The result is normalized and uploaded to an r8unorm texture.
  3. Lighting → a single TypeGPU fullscreen fragment pass samples camera and depth, computes radial falloff and a smoothstep occlusion test, and composites. Recorded into one GPUCommandEncoder per frame.

The lighting model

The light is treated as a real source in a pseudo-3D space, not a sprite pasted on top. Four things combine:

// 1. Pseudo-3D distance: screen offset plus a depth offset.
dist = sqrt(dot(dxy, dxy) + pow((disparity - lightDepth) * depthWeight, 2.0));

// 2. Inverse-square attenuation — an infinite tail, so light spills across the
//    room instead of stopping at a hard disc edge.
atten = 1.0 / (1.0 + pow(dist / lightRadius, 2.0));

// 3. Surfaces NEARER than the light have their camera-facing side turned away
//    from it, so they receive much less.
occlusion = smoothstep(lightDepth - softness, lightDepth + softness, disparity);
facing = 1.0 - 0.8 * occlusion;

// 4. Multiplicative, so the light REVEALS each surface's own colour and a dark
//    room genuinely brightens. Additive just washes a flat blob over it.
rgb = cam.rgb * (ambient + lightColor * lightIntensity * atten * facing * sourceGain);

sourceGain comes from bulb coverage: the depth is sampled at 37 points spread across the bulb's own disc, and the fraction that comes back occluded drives the dimming. Cover a quarter of the bulb and the room dims by about a quarter; cover all of it and you keep 10%, the way cupping a hand over a lamp leaks rather than blacking out. A single-point test would be all-or-nothing. The current figure is shown as bulb cov in the HUD.

smoothstep rather than a hard comparison throughout, because the model's depth boundaries are soft and a binary test aliases badly against them.

Why there is a depth-contrast control

Whatever is closest to the lens claims the top of the disparity range. A desk edge or keyboard in shot will squash everything else — including you — into the bottom few percent, which makes lightDepth impossible to tune. Two mitigations: the normalization window is a 2%–90% percentile (not min/max), and depth contrast applies a gamma that expands the range you actually occupy.

Measured performance

On an RTX 5080 in Chrome:

Stage Time
Render (lighting + composite) ~0.15–0.20 ms
Depth inference ~35–41 ms
CPU tensor convert + normalize ~0.7–1.1 ms
Depth texture upload ~0.03 ms
Display 180–280 fps

Display rate stays high because inference is decoupled — the last depth texture is reused until a new one arrives. But inference itself is ~5× slower than the 8 ms target. Transformers.js runs the ONNX graph through onnxruntime-web, which reports that some nodes fall back to CPU. There is also an unavoidable GPU→CPU→GPU round trip, since the tensor comes back on the CPU.

Closing that gap is Phase 6: reimplementing the model as hand-written TypeGPU compute kernels (~250 dispatches) so the depth buffer never leaves the GPU and joins the same command encoder as the lighting pass. That is a separate, multi-week project — build a per-layer numerical diff harness against ONNXRuntime first.

Layout

File Role
src/main.ts Orchestration, GUI, stats, input
src/renderer.ts TypeGPU root, shaders, bind groups, command encoder
src/depth.ts Model loading, inference, normalization
src/camera.ts getUserMedia + device enumeration
src/params.ts Shared uniform schema and presets
NOTES.md Debugging log — read this before deep-diving a bug

Notes

  • Depth Anything V2 Small only. Base and Large are CC-BY-NC.
  • A GPUExternalTexture expires at the end of the frame's task: the import and the bind group referencing it are both rebuilt every frame. Never cache them.

Author

Built by @AkbarSheikh-debug.

License

MIT — use it, fork it, ship it. The one condition is that the copyright notice in LICENSE stays in copies and substantial reuses of this code; that's not a suggestion, it's the actual license term. If you build on this, a link back or a mention is appreciated but not required beyond that.

About

No description, website, or topics provided.

Resources

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages