Skip to content

Pre-fork workers so requests don't wait on fork - #67

Merged
flavorjones merged 1 commit into
masterfrom
prefork-warmup-543
Sep 28, 2026
Merged

flavorjones merged 1 commit into
masterfrom
prefork-warmup-543

Conversation

@flavorjones

@flavorjones flavorjones commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Motivation

With the default max_requests_per_worker: 1, a cell's supervisor forks one worker process per request. On master, the supervisor forks that worker only when a request arrives. Each request waits for the fork and for the new worker to start. Only then does the operation run. ADR 0001 measured the fork alone at 2.8ms. It also noted that pre-forking would recover that time.

Details

  • Supervisor#run forks a worker into each free slot before it starts its loop. The first requests find workers that are ready.
  • When Supervisor#reap reaps a worker that served a request, it forks a replacement immediately. The next request does not wait for a fork.
  • Supervisor#reap does not replace a worker that exited without serving a request. This guards against a worker that crashes during its startup, for example after a bad deploy. A replacement would crash too, and the supervisor would fork again and again, with no requests arriving. Instead, the next request forks its own worker, as on master.
  • Process.warmup runs one time, at boot, in Supervisor#boot. It runs after the operations load and before the first fork. It compacts the supervisor's Ruby heap. Each worker then copies fewer inherited pages when its garbage collector writes to them.

Benchmark

Each run boots a cell with concurrency 2. The run sends Active Storage transforms through activestorage-hotcell-client, one at a time. It pauses between requests to simulate idle time. Each table cell shows the range of the median across runs.

The cell in ruby:3.4-slim (Ruby 3.4.11, libvips 8.16) with the accessory's flags:

master this branch this branch without Process.warmup
60×40 PNG → 30×30, 1s pauses 23.0–24.8ms 16.8–18.5ms 19.5–20.9ms
60×40 PNG → 30×30, 100ms pauses 22.5–23.3ms 10.2–17.4ms 12.0–19.7ms
3000×2000 JPEG → 800×600, 1s pauses 92.6–103.4ms 93.4–100.5ms 98.4–102.6ms
3000×2000 JPEG → 800×600, 100ms pauses 94.1–102.1ms 88.9–100.8ms 84.0–100.6ms

On the host (Ruby 4.0.3, libvips 8.18):

master this branch this branch without Process.warmup
60×40 PNG, 1s pauses 34.4–35.1ms 20.0–21.1ms 25.5–31.7ms
60×40 PNG, 100ms pauses 28.9–29.4ms 20.0–21.7ms 21.6–25.0ms

On the small image, pre-forking and Process.warmup each save time. With 1s pauses, the ranges do not overlap. On the large image, the image work takes most of the time. The differences are within the run-to-run variation.

Additional information

This PR also fixes a client race. Pre-forking made the race common. A client connects to the cell's socket. Then the client writes its request. When the cell is full, the supervisor accepts the connection and writes capacity. Then the supervisor closes the connection without reading the request (Supervisor#refuse). If the close happens before the client writes, the write fails with EPIPE. The client then returned unavailable. It did not read the capacity answer that was already on its socket. Now Transport::Socket reads the answer when its write fails with EPIPE or ECONNRESET. The new test in classification_test.rb sends a request that is larger than the socket buffer. The test server answers and closes without reading. So the write fails every time.

Copilot AI balanced review requested due to automatic review settings September 28, 2026 16:31

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@flavorjones flavorjones changed the title Pre-fork workers and warm up the supervisor's heap Pre-fork workers so requests don't wait on fork Sep 28, 2026
With `max_requests_per_worker: 1`, the supervisor forked a worker only when
a request arrived. Each request waited for the fork and for the worker to
start. Fork a worker into each free slot at boot. When a worker that served
a request is reaped, fork a replacement immediately. Run `Process.warmup`
one time, at boot, before the first fork.

This change made a client race common. A full cell writes `capacity` and
closes the connection without reading the request. If the close happened
before the client finished its write, the write failed with EPIPE. The
client then returned `unavailable` and did not read the `capacity` answer.
Read the answer when the write fails with EPIPE or ECONNRESET.

Median per request, sequential transforms 1s apart, the cell in
ruby:3.4-slim with the accessory's flags, concurrency 2:

  60x40 png thumbnail        master 23.8ms  this 17.1ms
  3000x2000 jpg to 800x600   master 99.9ms  this 99.3ms  (within noise)
@flavorjones
flavorjones merged commit cc8dc2e into master Sep 28, 2026
16 checks passed
@flavorjones
flavorjones deleted the prefork-warmup-543 branch September 28, 2026 20:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants