Skip to content

[Benchmark] Large-schema compile-time & binary-size comparison (600/1500 elements) deferred from #49 #55

Description

@nth-bailey

Context

The #49 perfect-hash element dispatch work shipped complete throughput and hardware-counter results for all four benchmark tiers (16/120/600/1500 elements), but the compile-time & binary-size half of the analysis could only be measured for a small schema: on the 7.7 GiB reference host (WSL2, ~5.4 GiB available), compiling the generated consumer crates at the 600- and 1500-element tiers exceeds the portable 60% RAM cap (scripts/memcap.sh, ≈3.2 GiB) for a single rustc invocation.

This is the follow-up issue referenced in docs/benchmarks/rust-phf-dispatch.md § "Compilation time & binary size" and in .agents/skills/polyxml-codegen-workflow/SKILL.md.

What was encountered (for the record)

  • Both variants die at both tiers. In-cgroup anon-RSS at the OOM kill was 3.19–3.21 GiB — right at the cap — for match and phf output at 600 and 1500 elements. Memory scales with field count (a single decode_xml spans 5423 lines at 1500 elements), not with dispatch strategy, so no match-vs-phf compile-memory conclusion can be drawn yet.
  • Uncapped runs are unsafe on 8-GiB-class hosts. Pre-memcap attempts swap-thrashed the box into freezing (two WSL restarts). Every capped attempt since was kernel-OOM-killed inside its own cgroup (exit 143) with the host untouched — do not retry uncapped; raise the cap deliberately or use a bigger host.
  • Small-schema baseline (measured, in the doc): match rebuild 2.51–2.54 s, .text=2590640 .rodata=367768; phf rebuild 2.53–2.55 s, .text=2591728 .rodata=367912 — Δ +1088 B .text, +144 B .rodata (+0.04 %), cold phf dep build +0.2 s one-time.

What to do

On a host where one rustc may safely use ≥ 4 GiB (16 GiB+ RAM, or POLYXML_MEMCAP_PCT raised deliberately):

  • Generate the 600-element schema tier (scripts/gen_tag_dispatch_fixtures.py) in both modes (default match, --feature phf + phf dep) and build the standalone consumer crate under scripts/memcap.sh: record cold build and touch-only rebuild times.
  • Same for the 1500-element tier.
  • Capture section sizes with size/size -A (.text, .rodata, .data, .bss) for all four crates.
  • Replace the "not measured — deferred" section of docs/benchmarks/rust-phf-dispatch.md with the full table and interpretation.
  • Re-check the ≥ ~900-element auto-enable recommendation if compile cost or .rodata growth at scale diverges materially from the small-schema delta (phf tables should stay tiny vs. giant match bodies — verify).

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    codegenPolyXML polyglot code generationenhancementNew feature or requesttarget:rustRust code generator target

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions