perf: mount new components 11-20x cheaper with an optional compiled core - #59
Draft
maartenbreddels wants to merge 16 commits into
Draft
maartenbreddels wants to merge 16 commits into
maartenbreddels wants to merge 16 commits into
Conversation
Every render of an element wrote two attributes: the frozen-key flag and the render count, always next to each other. The flag is now a property (rendered at least once), so a render writes one attribute. For a mount this is one write less per element; it matters more once the mount is compiled, where the remaining work per element is a few field writes. Behavior: an element whose render raised a duplicate-key KeyError (the pass is aborted) is no longer frozen, so .key() on it works again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The generated element factories call ComponentWidget(widget=cls) for every element they make (ipyvuetify's generated components, which solara uses for every v.* element, do too). The lookup of the shared instance in a WeakValueDictionary is a Python-level method call; reading a class attribute is a fraction of that. A subclass inherits the attribute, so the cached instance is only used when its widget is that exact class. The class -> instance -> class cycle is freed by gc like any class, so a widget class made at runtime is still not kept alive (test_dynamic_widget_class_is_freed). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every generated factory (reacton.ipywidgets, bqplot, ipycanvas) looked up the shared ComponentWidget of its widget class for every element it made. The component is now made once, when the module is imported, and the factory passes it to the element directly: one call less per element. The generator template does the same, so ipyvuetify's components (solara's v.* elements) get it when they are generated again with this reacton. Behavior: the widget class is resolved when the module is imported, not at every call, so replacing a widget class in its module at runtime is no longer seen by the factory (solara only patches methods of widget classes). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The target for mounts is 10x less reacton time, and the prototypes showed that plain Python stops at about 9-10x, while the same code compiled with Cython in pure Python mode gets 15-30x. reacton/_fastcore.py is that module: plain Python that runs as it is (PyPy, no compiler, development), and a C extension when built with the optional setup_cython.py. All typing is in _fastcore.pxd, so the plain Python version does not pay for it. REACTON_CYTHON=0 forces the plain version when a compiled one is there (reacton/_fastcore_import.py). This first step moves the data of an element and the methods the renderers call for every element (the constructor with the container recording, key(), meta(), shared(), _arguments_changed()) into ElementBase and ValueElementBase, plus ContainerAdder, find_elements, the render thread-local and the component element call. reacton.core.Element and ValueElement are Python subclasses of those, so user code sees the same classes (subclassing, arbitrary attributes, weak references and Element[...] type hints keep working, see fastcore_test.py). Setting reacton.core.DEBUG also reaches the elements. Compiled, the element fields are C fields: the update paths, which stay Python, measured the same or a bit faster (the argument compare is compiled). Elements and their components can be pickled and copied again (the shared ComponentWidget needed its widget class in __new__). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fused mount of phase 2 still wrote everything the update paths use, for every element of a new subtree: string keys, used_keys, and the elements, widgets, children, element_to_widget and resolved_kwargs dicts of every component. Most mounted components are never updated, so this was most of the reacton time of a mount. The mount is now in _fastcore (so it can be compiled). A mounted component (_MountedContext) keeps only its element tree positionally: the widgets and child contexts in the order they were made. The dicts are made from that, with the same keys, when an update path, get_widget or state_get first uses them; a mounted subtree that goes away is removed and closed from it, in the same order as before. The hook containers are made when a hook first needs them. The rare cases (state set or an exception during the mount, shared elements, a widget that fails to be made, an explicit key that could match a positional key) undo the mounts of the pass into the render bookkeeping of the two phase walk, as before. Restored state (state_set) is mounted the same way. The lazily made containers are slots: CPython shares the key table of instance dicts for up to 30 names per class, and more names made every attribute access of the update paths slower. Tests: the dicts made from a mounted tree have the keys of the two phase walk, removal (with some components whose dicts were made) cleans up and closes in the same order in both renderers, nested restored state. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the mount in _fastcore, most of the time left was generic Python work per component and per element: attribute reads on the contexts and the render context, module attribute lookups, and the hooks and listener registration that were still interpreted. This types the mount in _fastcore.pxd (the elements are C fields, the walk is C functions), and moves the hot hooks there: use_state, use_ref, use_memo, use_effect and use_event do the work in _fastcore for the fast renderer (any other render context gets its own methods, as before), with the debug-log check made once per render. Setters of the fast renderer and the observers of on_<trait> listeners are small objects instead of closures; the component element and ComponentWidget(widget=cls) are made without interpreted code. Uncompiled, the same code runs as plain Python. Behavior changes: - A use_state setter of the fast renderer is a callable object (_fastcore._Setter), not a function; it behaves the same (the same checks, warnings and render). The default renderer keeps its closure. - The observer of an on_<trait> listener is a callable object (_fastcore._Listener), in both renderers. - reacton.core.use_state/use_ref/use_memo/use_effect and reacton.ipyvue.use_event are the _fastcore functions (compiled: not Python functions; inspect.signature and the docstrings work). - Assignments to reacton.core._default_container and _component_context_manager_classes are passed on to _fastcore. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
solara appends its context manager class to reacton.core._component_context_manager_classes after importing reacton, and assigns reacton.core._default_container; the compiled mount reads both as its own module globals. The new test checks both, for a first mount and for a later update, and runs in both modes. The .pxd says these must stay plain module globals. The removal of a mounted component looked up the widgets dict (a Python call) for every component, only needed when a widget has orphans. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Profiles of the compiled mount (macOS sample) showed generic attribute reads and writes on Python objects as the main cost left besides the calls into user code. Per widget element the mount read the rerender flag of the render context and three facts on the ComponentWidget; the flag only changes in component bodies, so it is now checked after each body, and the facts of a widget class are one typed object on its ComponentWidget. The removal of a mounted component switched the render context's current component even when nothing could read it, and the render pass computed the stale keys of the root with a set difference even when the sizes show there are none. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fast renderer's hooks in _fastcore were a copy of the methods of reacton.core._RenderContext, and the mount had its own copy of the body call with the implicit container. Both renderers now use the one implementation in _fastcore (the default renderer with its own setter closure, made by make_setter), which also compiles the hooks and the body call of the default renderer when reacton is built with Cython. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ADME Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
solara.tasks does not warn about task() inside use_memo: it looks on the stack for a function named use_memo in a reacton module. Compiled functions have no frame, so the compiled fast renderer made solara's task_test warn. A thin Python use_memo keeps that frame; it costs one Python call per use_memo, only in the compiled mode. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
In the interleaved run of all 15 scenarios the compiled mount was only 7.6-9.8x on rows300, deep100 and page: about 2x the compiled floor, and more so in a cold cache (the benchmark collects garbage before each sample, so the first touch of every function and object counts). The fixed cost of a render was mostly Python frames and generic walks for the root: render_fixed, the render context __init__ chain, reading REACTON_FAST through os.environ, render() with its _render, _reconsolidate and the other root walks. render_fixed and the first render of a fast render context now run compiled (render_first); when more passes are needed (state set or an exception during the mount, an effect that sets state) the loop of render() takes over, now a method of its own. Per component, the mount did work the floor does not: a Python frame for use_context and for each Effect, a list per body for the implicit container, a list per setter, an isinstance of a Python class that fails (slow), a dict per element that is rarely used, and the registration of every on_<trait> observer in Element._callback_wrappers. The observers now stay in the node until the dicts are made (materialize registers them), and removing the node unobserves them. render() keeps the thread ident instead of the Thread object for its recursion check, and the reconsolidating flag tells an aborted render pass from a failed reconciliation (one attribute less on the render context: its instance dict stays within CPython's 30 shared keys). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
After the first batch, rows300, page and subtree_swap still spent more per unit than the compiled floor. Per widget: a bound method for the lookup of the element class flags (an untyped dict global), and the trait check of every kwarg, needed only when a value can be a listener (a callable or None). Per removed mounted component: two exception lists set on the context (and its dict grown for them) although almost no removal raises, an isinstance of a Python class that fails for every widget node, and a Python frame to stop the widget recording. The exceptions of a removal are now collected in local lists and bubble up as before (a test checks that a use_exception component above a removed child gets the exception of its effect cleanup, in both renderers). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ount deep100 was still at about 10x in the interleaved run: its mount walk now costs the same as the compiled floor's, so what is left is the fixed cost of a render, in a cold cache. Per render context, a deque for the rerender reasons (a 64 slot block from the system allocator, and keyword parsing) and a ThreadSafeCounter with its own lock were made although most render contexts never need them: both are made on first use now (the counter under a module lock, as two threads can start a batch). The logging check of the first render reads the cache dict of the logger instead of a Python call. After a mount, finish_mount visited every new context to look for effects; in a cold cache that is a miss per context. The mount now records only the contexts that have effects (it reads that while the context is still in the cache); the rare exception of an effect goes up through the new contexts with a walk over their nodes (a test checks a use_exception component above a new subtree, in both renderers). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README explained the mount path, but not that render_fixed and the first render of a new render context now skip the walks of render(): that is most of what changed in the fixed cost of a mount. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Python 3.7 raises "annotated name can't be global" when a module level annotation comes after a function that declares the name global; newer Pythons accept it. The installation check of CI runs on 3.7, so the module failed to import there. The declaration now sits with the other names that _register_hooks sets. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on the update PR. This makes mounts (first render, new list items, component type swaps) in the fast renderer 11-20x cheaper than
master, with an optional Cython-compiled core. Without Cython, the same code runs as plain Python; mounts are then 5-7x cheaper.Why
After the update work, mounts were only 3-4x cheaper, and a floor analysis showed why. An ideal pure-Python renderer with the same element and hook API only reaches about 9-10x, because what is left per element is Python call and object overhead. Compiling the hot path removes that.
What
reacton/_fastcore.py: the element building blocks, the mount, the hooks and the listeners, in one module written for Cython's pure-Python mode._fastcore.pxd, so the plain module does not pay for it.REACTON_CYTHON=0forces the plain module when a compiled one is present.render_first), and the render context's rarely used parts are created on first use.ComponentWidgetonce, and it is kept on the widget class.setup_cython.pyis optional:python setup_cython.py build_ext --inplace.pip install -e .without Cython keeps working, and the wheel stayspy3-none-any.Numbers
reacton's own time, fast renderer, against
master, same interleaved run:Behavior changes
_reacton_component_widgeton widget classes, including third-party ones._key_frozenis derived from the render count.use_contexteffect are small callable objects, not closures.vars(el)does not show themrender_fixedand the hooks have no Python frames, andgetsource()does not work on themuse_memokeeps a thin Python wrapper, because solara'staskslooks for its frameNot done yet
Tests
In both modes (compiled and
REACTON_CYTHON=0):pytest reacton/: 250 and 253 passed, with the default and the fast renderermasterin all four mode/renderer combinations🤖 Generated with Claude Code