Skip to content

perf: mount new components 11-20x cheaper with an optional compiled core - #59

Draft
maartenbreddels wants to merge 16 commits into
perf/render-updatesfrom
perf/render-10x
Draft

maartenbreddels wants to merge 16 commits into
perf/render-updatesfrom
perf/render-10x

Conversation

@maartenbreddels

Copy link
Copy Markdown
Contributor

Summary

Stacked on the update PR. This makes mounts (first render, new list items, component type swaps) in the fast renderer 11-20x cheaper than master, with an optional Cython-compiled core. Without Cython, the same code runs as plain Python; mounts are then 5-7x cheaper.

Why

After the update work, mounts were only 3-4x cheaper, and a floor analysis showed why. An ideal pure-Python renderer with the same element and hook API only reaches about 9-10x, because what is left per element is Python call and object overhead. Compiling the hot path removes that.

What

  • reacton/_fastcore.py: the element building blocks, the mount, the hooks and the listeners, in one module written for Cython's pure-Python mode.
    • All typing is in _fastcore.pxd, so the plain module does not pay for it.
    • REACTON_CYTHON=0 forces the plain module when a compiled one is present.
  • A new mount data model: a mounted component keeps only its positional node list. The dicts the update paths need are built on first use, with the same keys. The rare cases undo the mount into the two-phase bookkeeping, as before.
  • The first render is compiled too (render_first), and the render context's rarely used parts are created on first use.
  • The generated factories (ipywidgets, bqplot, ipycanvas) create their ComponentWidget once, and it is kept on the widget class.
  • Both renderers share one copy of the hooks.
  • setup_cython.py is optional: python setup_cython.py build_ext --inplace. pip install -e . without Cython keeps working, and the wheel stays py3-none-any.

Numbers

reacton's own time, fast renderer, against master, same interleaved run:

scenario compiled plain Python
mount 1024 buttons 19.9x 7.2x
mount 300 rows 14.1x 6.5x
mount 100 deep 11.2x 5.1x
mount page (227 components) 11.7x 5.2x
subtree swap 12.7x 6.1x
updates 10.8-61x 9.7-58x

Behavior changes

  • reacton sets _reacton_component_widget on widget classes, including third-party ones.
  • The generated factories resolve their widget class at import.
  • _key_frozen is derived from the render count.
  • Setters, listeners and the use_context effect are small callable objects, not closures.
  • Compiled mode:
    • element fields are C fields, so vars(el) does not show them
    • render_fixed and the hooks have no Python frames, and getsource() does not work on them
    • use_memo keeps a thin Python wrapper, because solara's tasks looks for its frame

Not done yet

  • CI does not build or test the compiled mode. Wheels would need a hatch build hook, cibuildwheel, and Cython pinned to 3.1.x, with the pure wheel as a fallback.
  • The plain-Python mount path can still be tuned: fewer helper calls, simpler element recording, and fewer hook objects. An estimate is about 8-12x without a compiler.

Tests

In both modes (compiled and REACTON_CYTHON=0):

  • pytest reacton/: 250 and 253 passed, with the default and the fast renderer
  • a tree check across implementations, and fuzz 300 x 40
  • the solara 1.62 unit suite gives the same failures as master in all four mode/renderer combinations

🤖 Generated with Claude Code

maartenbreddels and others added 16 commits September 26, 2026 03:31
Every render of an element wrote two attributes: the frozen-key flag and
the render count, always next to each other. The flag is now a property
(rendered at least once), so a render writes one attribute. For a mount
this is one write less per element; it matters more once the mount is
compiled, where the remaining work per element is a few field writes.

Behavior: an element whose render raised a duplicate-key KeyError (the
pass is aborted) is no longer frozen, so .key() on it works again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The generated element factories call ComponentWidget(widget=cls) for
every element they make (ipyvuetify's generated components, which solara
uses for every v.* element, do too). The lookup of the shared instance
in a WeakValueDictionary is a Python-level method call; reading a class
attribute is a fraction of that. A subclass inherits the attribute, so
the cached instance is only used when its widget is that exact class.
The class -> instance -> class cycle is freed by gc like any class, so a
widget class made at runtime is still not kept alive
(test_dynamic_widget_class_is_freed).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every generated factory (reacton.ipywidgets, bqplot, ipycanvas) looked up
the shared ComponentWidget of its widget class for every element it
made. The component is now made once, when the module is imported, and
the factory passes it to the element directly: one call less per
element. The generator template does the same, so ipyvuetify's
components (solara's v.* elements) get it when they are generated again
with this reacton.

Behavior: the widget class is resolved when the module is imported, not
at every call, so replacing a widget class in its module at runtime is no
longer seen by the factory (solara only patches methods of widget
classes).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The target for mounts is 10x less reacton time, and the prototypes showed
that plain Python stops at about 9-10x, while the same code compiled with
Cython in pure Python mode gets 15-30x. reacton/_fastcore.py is that
module: plain Python that runs as it is (PyPy, no compiler, development),
and a C extension when built with the optional setup_cython.py. All
typing is in _fastcore.pxd, so the plain Python version does not pay for
it. REACTON_CYTHON=0 forces the plain version when a compiled one is
there (reacton/_fastcore_import.py).

This first step moves the data of an element and the methods the
renderers call for every element (the constructor with the container
recording, key(), meta(), shared(), _arguments_changed()) into
ElementBase and ValueElementBase, plus ContainerAdder, find_elements, the
render thread-local and the component element call. reacton.core.Element
and ValueElement are Python subclasses of those, so user code sees the
same classes (subclassing, arbitrary attributes, weak references and
Element[...] type hints keep working, see fastcore_test.py). Setting
reacton.core.DEBUG also reaches the elements. Compiled, the element
fields are C fields: the update paths, which stay Python, measured the
same or a bit faster (the argument compare is compiled).

Elements and their components can be pickled and copied again (the
shared ComponentWidget needed its widget class in __new__).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fused mount of phase 2 still wrote everything the update paths use,
for every element of a new subtree: string keys, used_keys, and the
elements, widgets, children, element_to_widget and resolved_kwargs dicts
of every component. Most mounted components are never updated, so this
was most of the reacton time of a mount.

The mount is now in _fastcore (so it can be compiled). A mounted
component (_MountedContext) keeps only its element tree positionally:
the widgets and child contexts in the order they were made. The dicts
are made from that, with the same keys, when an update path, get_widget
or state_get first uses them; a mounted subtree that goes away is
removed and closed from it, in the same order as before. The hook
containers are made when a hook first needs them. The rare cases (state
set or an exception during the mount, shared elements, a widget that
fails to be made, an explicit key that could match a positional key)
undo the mounts of the pass into the render bookkeeping of the two phase
walk, as before. Restored state (state_set) is mounted the same way.

The lazily made containers are slots: CPython shares the key table of
instance dicts for up to 30 names per class, and more names made every
attribute access of the update paths slower.

Tests: the dicts made from a mounted tree have the keys of the two phase
walk, removal (with some components whose dicts were made) cleans up and
closes in the same order in both renderers, nested restored state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the mount in _fastcore, most of the time left was generic Python
work per component and per element: attribute reads on the contexts and
the render context, module attribute lookups, and the hooks and listener
registration that were still interpreted. This types the mount in
_fastcore.pxd (the elements are C fields, the walk is C functions), and
moves the hot hooks there: use_state, use_ref, use_memo, use_effect and
use_event do the work in _fastcore for the fast renderer (any other
render context gets its own methods, as before), with the debug-log
check made once per render. Setters of the fast renderer and the
observers of on_<trait> listeners are small objects instead of closures;
the component element and ComponentWidget(widget=cls) are made without
interpreted code. Uncompiled, the same code runs as plain Python.

Behavior changes:
- A use_state setter of the fast renderer is a callable object
  (_fastcore._Setter), not a function; it behaves the same (the same
  checks, warnings and render). The default renderer keeps its closure.
- The observer of an on_<trait> listener is a callable object
  (_fastcore._Listener), in both renderers.
- reacton.core.use_state/use_ref/use_memo/use_effect and
  reacton.ipyvue.use_event are the _fastcore functions (compiled: not
  Python functions; inspect.signature and the docstrings work).
- Assignments to reacton.core._default_container and
  _component_context_manager_classes are passed on to _fastcore.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
solara appends its context manager class to
reacton.core._component_context_manager_classes after importing reacton,
and assigns reacton.core._default_container; the compiled mount reads
both as its own module globals. The new test checks both, for a first
mount and for a later update, and runs in both modes. The .pxd says
these must stay plain module globals.

The removal of a mounted component looked up the widgets dict (a Python
call) for every component, only needed when a widget has orphans.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Profiles of the compiled mount (macOS sample) showed generic attribute
reads and writes on Python objects as the main cost left besides the
calls into user code. Per widget element the mount read the rerender
flag of the render context and three facts on the ComponentWidget; the
flag only changes in component bodies, so it is now checked after each
body, and the facts of a widget class are one typed object on its
ComponentWidget. The removal of a mounted component switched the render
context's current component even when nothing could read it, and the
render pass computed the stale keys of the root with a set difference
even when the sizes show there are none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fast renderer's hooks in _fastcore were a copy of the methods of
reacton.core._RenderContext, and the mount had its own copy of the body
call with the implicit container. Both renderers now use the one
implementation in _fastcore (the default renderer with its own setter
closure, made by make_setter), which also compiles the hooks and the
body call of the default renderer when reacton is built with Cython.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ADME

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
solara.tasks does not warn about task() inside use_memo: it looks on the
stack for a function named use_memo in a reacton module. Compiled
functions have no frame, so the compiled fast renderer made solara's
task_test warn. A thin Python use_memo keeps that frame; it costs one
Python call per use_memo, only in the compiled mode.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
In the interleaved run of all 15 scenarios the compiled mount was only
7.6-9.8x on rows300, deep100 and page: about 2x the compiled floor,
and more so in a cold cache (the benchmark collects garbage before each
sample, so the first touch of every function and object counts).

The fixed cost of a render was mostly Python frames and generic walks
for the root: render_fixed, the render context __init__ chain, reading
REACTON_FAST through os.environ, render() with its _render,
_reconsolidate and the other root walks. render_fixed and the first
render of a fast render context now run compiled (render_first); when
more passes are needed (state set or an exception during the mount, an
effect that sets state) the loop of render() takes over, now a method
of its own.

Per component, the mount did work the floor does not: a Python frame
for use_context and for each Effect, a list per body for the implicit
container, a list per setter, an isinstance of a Python class that
fails (slow), a dict per element that is rarely used, and the
registration of every on_<trait> observer in Element._callback_wrappers.
The observers now stay in the node until the dicts are made
(materialize registers them), and removing the node unobserves them.

render() keeps the thread ident instead of the Thread object for its
recursion check, and the reconsolidating flag tells an aborted render
pass from a failed reconciliation (one attribute less on the render
context: its instance dict stays within CPython's 30 shared keys).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
After the first batch, rows300, page and subtree_swap still spent more
per unit than the compiled floor. Per widget: a bound method for the
lookup of the element class flags (an untyped dict global), and the
trait check of every kwarg, needed only when a value can be a listener
(a callable or None). Per removed mounted component: two exception
lists set on the context (and its dict grown for them) although almost
no removal raises, an isinstance of a Python class that fails for every
widget node, and a Python frame to stop the widget recording.

The exceptions of a removal are now collected in local lists and
bubble up as before (a test checks that a use_exception component above
a removed child gets the exception of its effect cleanup, in both
renderers).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ount

deep100 was still at about 10x in the interleaved run: its mount walk
now costs the same as the compiled floor's, so what is left is the
fixed cost of a render, in a cold cache. Per render context, a deque
for the rerender reasons (a 64 slot block from the system allocator,
and keyword parsing) and a ThreadSafeCounter with its own lock were
made although most render contexts never need them: both are made on
first use now (the counter under a module lock, as two threads can
start a batch). The logging check of the first render reads the cache
dict of the logger instead of a Python call.

After a mount, finish_mount visited every new context to look for
effects; in a cold cache that is a miss per context. The mount now
records only the contexts that have effects (it reads that while the
context is still in the cache); the rare exception of an effect goes up
through the new contexts with a walk over their nodes (a test checks a
use_exception component above a new subtree, in both renderers).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README explained the mount path, but not that render_fixed and the
first render of a new render context now skip the walks of render():
that is most of what changed in the fixed cost of a mount.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Python 3.7 raises "annotated name can't be global" when a module level
annotation comes after a function that declares the name global; newer
Pythons accept it. The installation check of CI runs on 3.7, so the
module failed to import there. The declaration now sits with the other
names that _register_hooks sets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant