Skip to content

fix(render): stop requesting media data when a pump runs dry - #43

Open
bilipp wants to merge 1 commit into
mainfrom
fix/render-pump-starvation-spin
Open

bilipp wants to merge 1 commit into
mainfrom
fix/render-pump-starvation-spin

Conversation

@bilipp

@bilipp bilipp commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Fixes the SystemRenderer pumps pegging two CPU cores whenever their decode channels are empty, which is most of the time on live TV.

What happens

A customer reported crashes while zapping on LumeEngine. Their tvOS 27 cpu_resource reports from Lume 2.2.0 (25) on an Apple TV 4K (3rd gen) show:

  • 69% and 91% CPU averaged over 130 s and 99 s.
  • Two hot threads at userInteractive.
  • Every sample under AVFoundation's requestMediaDataWhenReady callback, calling into LumeEngine.

Why

requestMediaDataWhenReady is not edge-triggered. As long as the renderer stays ready, AVFoundation calls the block again the moment it returns. A standalone check on macOS ran a starved block about 2.6M (video) and 3.2M (audio) times per second, using two full cores.

The pumps returned as soon as tryReceive() came back empty, with the renderer still ready. The comment on pumpVideo assumed the block only re-fires on readiness transitions. On live TV, frames arrive in real time and the renderer never fills, so both pumps spin until the stream ends. During a zap the channels are empty while the new stream buffers, which is the worst case.

Fix

  • Starved pump: calls stopRequestingMediaData(), and its retry re-arms requestMediaDataWhenReady after 10 ms. The retry used to call the pump directly after 30 ms; the 10 ms is the only latency a frame landing in an empty channel picks up.
  • Pump with no lane (detached or shut down): stops requesting instead of returning, which was the same spin.
  • Arming order: requests are always armed on the lane's own queue, the queue the pump stops on. That orders the two, so an audio track switch (attach(audio: nil) then attach(audio: frames)) can't have its fresh request cancelled by a pump still running for the old lane. Registering twice just replaces the block.
  • pumpCounts: a new internal, lock-guarded read of how often each pump ran.

Tests

New SystemRendererTests:

  • starvedPumpsIdleOnTheirRetry: attached lanes with empty channels must stay under 250 pump runs in 0.5 s. Before the fix: 271,717 (video) and 286,723 (audio).
  • starvedVideoPumpPicksUpALateFrame: a frame sent after the pump went quiet still reaches the renderer (PLAN.md §3.3, no stuck spinner).

swift test passes: 106 tests in 17 suites. The engine also compiled for the tvOS Simulator in a Lume Release build.

Not verified: a real live stream on an Apple TV, to confirm CPU on the device.

Shipping

The App Store build picks this up from the local ../LumeEngine checkout. The sideload workflow needs its ENGINE_REF pin bumped once this is tagged.

`requestMediaDataWhenReady` is not edge-triggered. For as long as the
renderer is ready, AVFoundation calls the block again the moment it
returns. The pumps returned on an empty channel with the renderer still
ready, so they ran millions of times a second: two cores pegged at
userInteractive. Live TV is in that state almost all the time, because
frames arrive in real time and the renderer never fills. Customer tvOS
cpu_resource reports while zapping show 69-91% CPU, all of it in this loop.

- A starved pump stops requesting, and its retry re-arms the request after
  10 ms (was 30 ms, which used to call the pump directly).
- A pump with no lane, detached or shut down, stops requesting instead of
  returning, the same spin.
- Requests are armed on the lane's own queue, the queue the pump stops on,
  so an audio track switch (detach, then re-attach) can't have its fresh
  request cancelled by a pump still running for the old lane. Registering
  twice just replaces the block.

SystemRendererTests counts pump runs against empty channels (about 272k per
lane in 0.5 s before, under 250 after) and checks that a frame arriving
after the pump went quiet still reaches the renderer.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant