Skip to content

[PERF] Cut per-VM allocations in the append VM - #21606

Closed
NullVoxPopuli wants to merge 2 commits into
emberjs:mainfrom
NullVoxPopuli:nvp/perf/append-vm-stacks
Closed

NullVoxPopuli wants to merge 2 commits into
emberjs:mainfrom
NullVoxPopuli:nvp/perf/append-vm-stacks

Conversation

@NullVoxPopuli

Copy link
Copy Markdown
Contributor

The Stacks class uses plain arrays instead of six StackImpl wrappers, and execute() runs the opcode loop directly instead of allocating a { done, value } result per instruction. A VM is constructed for every independently re-rendering block, so this shows up on {{#each}}-heavy pages.

🤖 Generated with Claude Code

Comment thread packages/@glimmer/runtime/lib/vm/append.ts Outdated
@NullVoxPopuli
NullVoxPopuli marked this pull request as draft September 9, 2026 22:53
@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

@NullVoxPopuli
NullVoxPopuli force-pushed the nvp/perf/append-vm-stacks branch from 023b930 to 6ae1674 Compare September 30, 2026 14:51
NullVoxPopuli and others added 2 commits September 30, 2026 15:19
Split out of nvp/simplify-some-vm-hot-paths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@NullVoxPopuli
NullVoxPopuli force-pushed the nvp/perf/append-vm-stacks branch from 6ae1674 to 01d7c35 Compare September 30, 2026 19:19
@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor

Tracerbench PDF for a run with gc() before the first measurement and a measured gc() at the end: tracerbench-report.pdf

With every sample starting from the same heap, this PR is a small net gain: Ember's script time is 2.1% lower in total, 4 to 7% lower on most renders, 5.7% lower on the second update, and 6.6% lower on the second select. The append slowdowns from my plain run are gone. Two costs remain: clearItems2 is 15% slower in script time (about 13 ms at 8x), and the app's first render is 14% slower (about 3 ms at 8x). The final gc() takes as long as on main, so this PR leaves no extra garbage behind.

phase main script ms (8x) script time full phase
all phases 9239 -2.1% no change
render 19 +14.2% slower no change
render1000Items1 604 no change -2.4%
clearItems1 113 no change -4.3%
render1000Items2 284 -4.1% no change
clearItems2 86 +15.1% slower +5.1% slower
render10000Items1 2197 -7.0% no change
render1000Items3 169 -6.5% -4.0%
updateEvery10thItem2 81 -5.7% no change
selectSecondRow1 121 -6.6% no change

Tracerbench's own report still marks selectSecondRow1 19.9% slower. With main-thread GC removed, that phase does not change, so a major GC still lands there in the middle of the run. The gc() calls reset the heap only at the start. The +112% on updateEvery10thItem2 from the plain run does not appear in this run.

How this was measured
  • ember.js pnpm bench settings: tracerbench compare on smoke-tests/benchmark-app, 8x CPU throttle, headless, fidelity 50.
  • Benchmark app change for this run, on control and experiment alike: gc() at the top of runBenchmark(), before the first mark, and a measured finalGc phase that calls gc() after clearItems4. Chrome runs with --js-flags=--expose-gc. Patch: benchmark-app-gc.patch
  • Control is main 153364bb9a. Experiment is this PR on top of 153364bb9a (merged locally when the branch is behind).
  • Chrome and tracerbench ran pinned to one CPU core. I also added --disable-background-networking,--disable-component-update to the Chrome flags and raised --sampleTimeout to 180 s.
  • The PDF is tracerbench's own report. The only change is that its results folder reads results/<label> instead of a local path.
  • Table values are the median of 50 per-round ratios. Tracerbench runs control and experiment back to back in each round, so both sides of a pair share the machine state. A value counts as a change when its bootstrap 95% CI excludes 0 and a Wilcoxon signed-rank test gives p < 0.05.
  • "Script time" runs from the phase's Start mark to the end of the task that holds it: the click plus Ember's render. It leaves out layout, paint, the benchmark's row checks and idle time. Main-thread GC inside that window is removed.
  • "Full phase" is tracerbench's own phase with main-thread GC removed.
  • An earlier main-vs-main run (without the gc() calls) had false hits of up to +5% on script time.
  • Raw results for every run: bench-reports branch.

@NullVoxPopuli
NullVoxPopuli deleted the nvp/perf/append-vm-stacks branch October 1, 2026 00:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants