From c2287c155a4b1a07d864e86d20273645f1a5b7d3 Mon Sep 17 00:00:00 2001 From: Pascal Severin Date: Fri, 25 Sep 2026 17:12:57 +0200 Subject: [PATCH 1/5] [M1-BENCH-01] laige-bench sim-tick suite: 10k-entity tick budget on both SimMath backends MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The PRD 8.1 simulation-tick workload as a laige-bench suite: 10k entities (2k dynamic bodies with Position2D + velocity, 8k bare), 60 Hz through the GameLoop on an exact synthetic clock, BenchMove + BenchHash (per-tick state hash) systems, 3000 measured ticks after 1000 warm-up, checked against sim_tick_avg (<=3.0 ms) and sim_tick_p99 (<=5.0 ms) on BOTH SimMath backends (ADR 0002). In debug non-sanitizer builds every measured tick additionally passes the engine's own G-R1 per-tick zero-allocation assertion (M1-ALLOC-01). Tool: new --math= option, repeatable --budget (one run gates every named budget against the same histogram), suite setup/teardown hooks, compiler-id fix (Clang before GCC — Clang 22's __VERSION__ is 'Clang 22.1.8' and the old order mislabeled clang builds). CI gate: laige_bench_sim_tick_{fpx16,fp32} ctest entries in every non-instrumented P0 job's full ctest (excluded from ASan/TSan trees — instrumentation inflates the tick cost, the m1-profiler-cost precedent); a budget FAIL exits 2 and fails the job (PRD 8.1 regression gate). Results (canonical Debug g++ 16.2.1, AMD Ryzen 9 7950X3D): fpx16_16 mean=0.59283 ms p99=0.618314 ms fp32_pinned mean=0.589047 ms p99=0.605951 ms Both backends PASS (5.1x / 8.1x inside the budgets). The run's final_hash is bit-identical across all non-instrumented local trees and both compilers, at both backends. budgets.json measured (worse of the two backends, first recorded values): sim_tick_avg 0.59283 ms / sim_tick_p99 0.618314 ms. Fourth baseline: docs/benchmarks/baselines/m1-sim-tick.md (full AGENTS 12 metadata, verbatim runs, cross-tree/cross-compiler table). Docs updated: benchmarks READMEs, budget_harness.md. --- budgets.json | 4 +- docs/api/budget_harness.md | 50 ++- docs/benchmarks/README.md | 14 +- docs/benchmarks/baselines/README.md | 16 +- docs/benchmarks/baselines/m1-sim-tick.md | 223 ++++++++++ roadmap/M1-heartbeat.md | 2 +- roadmap/README.md | 3 +- tools/bench/CMakeLists.txt | 58 ++- tools/bench/laige-bench.cpp | 513 ++++++++++++++++++++--- 9 files changed, 792 insertions(+), 91 deletions(-) create mode 100644 docs/benchmarks/baselines/m1-sim-tick.md diff --git a/budgets.json b/budgets.json index fa52532..222ad32 100644 --- a/budgets.json +++ b/budgets.json @@ -15,7 +15,7 @@ "metric": "mean", "unit": "ms", "target": 3.0, - "measured": 0, + "measured": 0.59283, "workload": "10k entities, 2k dynamic bodies (PRD 8.1)" }, { @@ -23,7 +23,7 @@ "metric": "p99", "unit": "ms", "target": 5.0, - "measured": 0, + "measured": 0.618314, "workload": "10k entities, 2k dynamic bodies (PRD 8.1)" }, { diff --git a/docs/api/budget_harness.md b/docs/api/budget_harness.md index 1d5ee8a..36a05ba 100644 --- a/docs/api/budget_harness.md +++ b/docs/api/budget_harness.md @@ -38,8 +38,34 @@ if (!r.passed) { /* print r.report; fail the CI run (PRD 8.1 policy) */ } ``` The end-to-end form is the tool: -`./build/bin/laige-bench --suite=synthetic --runs=1000 --warmup=100 -[--budget=]` — exit code 0 on pass, 2 on a failed budget check. + +```console +./build/bin/laige-bench --suite= [--runs=N] [--warmup=N] + [--budget=]... [--budgets=] + [--math=] + [--report=] [--list] +``` + +- `--suite` — required; `synthetic` (the M0 harness stand-in) or + `sim-tick` (the M1-BENCH-01 PRD §8.1 workload: 10k entities, 2k + dynamic bodies, 60 Hz, movement + per-tick state hash — its baseline: + `docs/benchmarks/baselines/m1-sim-tick.md`). `--list` prints the + suite table. +- `--budget` — **repeatable**: one run checks every named budget against + the same histogram (each entry's `metric` picks its statistic), + printing one report per entry. A duplicate name is a usage error. + The M1 gate form: `--budget=sim_tick_avg --budget=sim_tick_p99`. +- `--math` — the SimMath backend (ADR 0002) the `sim-tick` suite runs + on; `fixed_point_16_16` (the default) or `float_pinned_32`. Other + suites ignore it. +- `--budgets` — the budgets file (default: `$LAIGE_BUDGETS_PATH`, then + `budgets.json` in the working directory — the ctest gate entries set + the env var to the repo root). +- `--report` — append the printed report to a file. + +Exit code 0 on pass, 2 on a failed budget check (any one of the named +budgets failing fails the run — PRD 8.1 policy: a budget regression +fails CI), 1 on usage/load errors. ## `laige::Histogram` @@ -175,8 +201,10 @@ fields explicitly). Schema version **1**: the first baseline). - **All 15 PRD §8.1 targets** are present as named entries (sim tick, 50k-sprite scene, cold start, and the zone server each carry two - budgets). Until their subsystems exist (M1+) their `measured` values - stay 0 = not yet measured. + budgets). `sim_tick_avg`/`sim_tick_p99` carry their first recorded + values (M1-BENCH-01, `docs/benchmarks/baselines/m1-sim-tick.md` — + the worse of the two SimMath backends, canonical Debug tree); every + other entry stays 0 = not yet measured until its subsystem lands. - **Versioning:** an incompatible schema change bumps `version` and ships a migration note here; `loadBudgets` rejects every other version (never guesses). @@ -204,8 +232,12 @@ is smaller than the run and then claiming "all samples"; using ## Determinism The harness measures wall-clock durations: it is *not* deterministic -simulation state and plays no part in replay/lockstep (ARCH-010). The -synthetic `laige-bench --suite=synthetic` workload is deterministic by -construction (fixed LCG constants — Marsaglia 2003 — no RNG, no -allocation), which is what makes its baseline reproducible -(`docs/benchmarks/baselines/m0-synthetic.md`, M0-EXIT-01). +simulation state and plays no part in replay/lockstep (ARCH-010). Both +suites are deterministic by construction, which is what makes their +baselines reproducible: `synthetic` (fixed LCG constants — Marsaglia +2003 — no RNG, no allocation; `docs/benchmarks/baselines/m0-synthetic.md`, +M0-EXIT-01) and `sim-tick` (fixed seed `0x1F055EED`, deterministic +index-derived initial state, no randomness in the measured path — the +run's `final_hash` line is a bit-identical fingerprint across runs, +build types, and compilers per the ADR 0002 scope; +`docs/benchmarks/baselines/m1-sim-tick.md`, M1-BENCH-01). diff --git a/docs/benchmarks/README.md b/docs/benchmarks/README.md index 18b7acb..c378812 100644 --- a/docs/benchmarks/README.md +++ b/docs/benchmarks/README.md @@ -9,12 +9,14 @@ Benchmark methodology, baselines, results, and the regression policy map onto them, the baseline-file convention, the PRD §8.1 regression policy, and workload discipline. - [baselines/](baselines/README.md) — recorded baseline reports - (`-.md`). Every `budgets.json` entry still has - `measured: 0`: the two baselines recorded so far - (`baselines/m0-synthetic.md`, M0-EXIT-01, and - `baselines/m1-ecs-stress.md`, M1-ECS-07) measure harness and stress - workloads that update no budget. `baselines/m1-profiler-cost.md` - lands with M1-PROF-01. + (`-.md`). Four are recorded: + `baselines/m0-synthetic.md` (M0-EXIT-01), + `baselines/m1-ecs-stress.md` (M1-ECS-07), + `baselines/m1-profiler-cost.md` (M1-PROF-01), and + `baselines/m1-sim-tick.md` (M1-BENCH-01) — the first that updates a + `budgets.json` `measured` field (`sim_tick_avg`, `sim_tick_p99`); + every other entry still has `measured: 0` (its subsystem lands in a + later milestone). - [determinism-matrix.md](determinism-matrix.md) — the **determinism report** for the M1-SAMPLE-01 hello scenario (M1-DET-04): the committed per-tick hash baselines (both SimMath backends), the CI diff --git a/docs/benchmarks/baselines/README.md b/docs/benchmarks/baselines/README.md index 0617575..81c68ef 100644 --- a/docs/benchmarks/baselines/README.md +++ b/docs/benchmarks/baselines/README.md @@ -16,8 +16,7 @@ and the exact command + commit that produced it. Baseline files are **immutable**: superseding a baseline adds a new file and updates `measured` in `budgets.json` — it never edits an existing baseline. -Every `budgets.json` entry still has `measured: 0` (their subsystems land -in M1+). Recorded so far: +Recorded so far: - [m0-synthetic.md](m0-synthetic.md) (M0-EXIT-01, 2026-09-13) — the synthetic harness workload; proves the measurement pipeline end to @@ -27,3 +26,16 @@ in M1+). Recorded so far: of add/remove churn): no leaks (ASan), pool high-water stable, iteration within the documented cost, accounted ECS storage bytes. Not a `budgets.json` workload — no `measured` field updated. +- [m1-profiler-cost.md](m1-profiler-cost.md) (M1-PROF-01, + 2026-09-21) — the enabled-profiler cost gate (ON vs OFF on 10k-entity + ticks, bounded at 1%): measured +0.28% on the canonical Debug tree. + Not a `budgets.json` workload — the FR-11.1 counters are diagnostics, + not a budgeted subsystem. +- [m1-sim-tick.md](m1-sim-tick.md) (M1-BENCH-01, 2026-09-25) — the + PRD §8.1 simulation-tick workload (10k entities, 2k dynamic bodies, + 60 Hz, movement + per-tick state hash, both SimMath backends): + 0.59 ms mean / 0.62 ms p99 on the canonical Debug tree, both + backends PASS. **The first baseline to update `budgets.json` + `measured`** (`sim_tick_avg`, `sim_tick_p99` — the worse of the two + backends). Every other `budgets.json` entry still has + `measured: 0` (its subsystem lands in a later milestone). diff --git a/docs/benchmarks/baselines/m1-sim-tick.md b/docs/benchmarks/baselines/m1-sim-tick.md new file mode 100644 index 0000000..e88fadf --- /dev/null +++ b/docs/benchmarks/baselines/m1-sim-tick.md @@ -0,0 +1,223 @@ +# Baseline: `m1-sim-tick` — M1-BENCH-01 simulation-tick budget + +Recorded by **M1-BENCH-01** (2026-09-25). This is the **fourth** +baseline file; it is immutable (methodology §4 — superseding it later +adds a new file, it is never edited). + +## What this baseline measures + +The `sim-tick` suite of `laige-bench` (M1-BENCH-01): the PRD §8.1 +"Simulation tick" workload — **10 000 entities** (the scene budget at +100%), **2 000 of them dynamic bodies** — driven at **60 Hz** through +the engine's `GameLoop`, checked against **both** `sim_tick` budgets +(`sim_tick_avg`, `sim_tick_p99`) on **both** SimMath backends (ADR +0002: `fixed_point_16_16`, the default, and `float_pinned_32`). + +Workload shape (the suite is in `tools/bench/laige-bench.cpp`): + +- `World` capacity 10 000, deterministic, seed `0x1F055EED` (the + repo-wide test-seed convention, docs/testing.md §4; methodology §6). +- 10 000 entities; the first 2 000 are dynamic bodies carrying the + built-in `Position2D` plus the workload's velocity component + `BenchVel` (a `{Vec2 v}` component marked + `LAIGE_COMPONENT` + `LAIGE_DETERMINISM_SAFE` per backend). The + remaining 8 000 are bare entities (no components) — the static + scenery's M1 stand-in. +- Initial state is deterministic and index-derived (no RNG in the + measured path): the 2 000 dynamic bodies scatter over a 50×50 grid + span with per-tick velocities `vx ∈ -3..3`, `vy ∈ -5..5` world units + (max drift 20 025 units over the full 4 000-tick run — inside + `fpx16_16`'s ±32 768 Q16.16 range, so no saturating-overflow edge is + ever touched in the measured window). +- Two systems, registered in execution order (no `depends_on`): + - `BenchMove` (1 ms declared budget) — `pos += vel` over every + dynamic body, through the active backend's `SimMath::add` (the + pinned op surface, ADR 0002 — one correctly-rounded addition per + component, never fused); + - `BenchHash` (2 ms declared budget) — the per-tick deterministic + state hash (`World::stateHash`, M1-DET-03) over the live state, + labeled with the completed-tick count — the M1 determinism work + made part of the measured tick. +- The `GameLoop` (60 Hz) runs on an **exact synthetic clock**: each + measured iteration advances the injectable `nowNs` by exactly + `16 666 667` ns (one due tick — the exact integer due computation, + M1-LOOP-01) and runs one `frame()`. One measured sample is one + completed tick (beginFrame + the systems + the loop's bookkeeping). + The presentation snapshot (M1-LOOP-02) is not part of it — + presentation is separate from the authoritative tick by ARCH-009. +- **Warm-up:** 1 000 ticks discarded; **measured:** 3 000 ticks + (`--runs=3000 --warmup=1000`), every sample kept (histogram capacity + 3 000, no truncation). +- **Zero allocations:** in debug non-sanitizer builds, every measured + tick additionally passes the engine's own G-R1 per-tick + zero-allocation assertion (M1-ALLOC-01 — `GameLoop::runOneTick` arms + the process-wide allocation watch around the tick body). An + allocating tick aborts the run, so a PASS here is a zero-alloc- + asserted tick, not just a fast one. + +The measured sample is the **ECS-only slice** of the PRD §8.1 +workload: the 2 000 dynamic bodies move by kinematic integration +(movement + state hash). The PRD's remaining tick content (physics +bodies arrive with M3, presentation with the M4 render milestone) is +out of scope here — this is the slice M1 ships, measured at M1's +10 000-entity scene budget. + +The gate (roadmap M1-BENCH-01): `sim_tick_avg` mean ≤ 3.0 ms and +`sim_tick_p99` p99 ≤ 5.0 ms, **both backends**, enforced on CI by the +`laige_bench_sim_tick_fpx16` / `laige_bench_sim_tick_fp32` ctest +entries (one run checks both budgets via the repeatable +`--budget=` form; a failed check exits 2 and fails the CI job). The +entries are excluded from the sanitizer trees: instrumentation +inflates the tick's absolute cost (the m1-profiler-cost baseline +precedent — ASan+UBSan measures ≈1.6 ms/tick on this workload, 2.7× +the canonical 0.59 ms), so a sanitizer measurement would measure the +instrumentation, not the sim; those trees verify the workload's +safety properties instead (the leak-free ASan run, the race-free TSan +run — see Interpretation). + +**This is a `budgets.json` workload:** this baseline updates +`measured` for `sim_tick_avg` and `sim_tick_p99` (the M0 convention's +`0 = not yet measured` is retired by this measurement). The recorded +value is the **worse (max) of the two backends** on the canonical +Debug tree — the budget covers both backends (ADR 0002: both must +pass), and recording the worse keeps the 10% regression band +(docs/benchmarks/methodology.md §5) conservative. + +## AGENTS §12 metadata + +| # | Field | Value | +|---|---|---| +| 1 | Hardware | AMD Ryzen 9 7950X3D (16 cores / 32 threads), 64 GB RAM | +| 2 | OS | CachyOS (Arch-based Linux), kernel `7.2.6-1-cachyos`, x86_64 | +| 3 | Compiler and version | `g++ (GCC) 16.2.1 20260810` (canonical gate runs; the cross-compiler runs below use `Clang 22.1.8`) | +| 4 | Build type | `Debug` (canonical, `build/` tree) | +| 5 | Relevant flags | Engine policy (NFR-8.10): `-Wall -Werror -fno-exceptions -fno-rtti`; SimMath pinned set (ADR 0002) on the benchmark TU: `-ffp-contract=off -fno-associative-math`. No sanitizers (canonical tree). | +| 6 | Dataset / workload | `sim-tick` — 10 000 entities at the 100% scene budget, 2 000 dynamic bodies (`Position2D` + `BenchVel`), 8 000 bare entities; systems `BenchMove` (1 ms budget, `pos += vel` via `SimMath::add`) and `BenchHash` (2 ms budget, per-tick `World::stateHash`); `GameLoop` 60 Hz, 1 tick/frame, exact synthetic clock (`16 666 667` ns/frame); both SimMath backends | +| 7 | Warm-up | 1 000 ticks discarded | +| 8 | Sample count | `n=3000` tick samples per backend (histogram capacity 3000, no truncation) | +| 9 | Summary statistics | `fpx16_16`: mean=0.59283 p50=0.591775 p95=0.604027 **p99=0.618314** max=0.679561 ms · `fp32_pinned`: mean=0.589047 p50=0.588107 p95=0.599378 **p99=0.605951** max=0.717022 ms (min: 0.580593 / 0.578499) | +| 10 | Before / after | `before=0` (M0 convention — not yet measured) · `after`: sim_tick_avg (mean) `fpx16=0.59283`, `fp32=0.589047` → **recorded 0.59283** (worse of backends); sim_tick_p99 (p99) `fpx16=0.618314`, `fp32=0.605951` → **recorded 0.618314** (worse of backends) · `target`: mean ≤ 3.0 ms, p99 ≤ 5.0 ms — both backends **PASS** (5.1× / 8.1× inside the gates) | + +## Verbatim run output + +### Run — canonical tree (Debug, g++), both backends, budget-checked + +Command (run from the repository root; `budgets.json` resolved via +`LAIGE_BUDGETS_PATH`): + +```console +$ LAIGE_BUDGETS_PATH=$PWD/budgets.json ./build/bin/laige-bench \ + --suite=sim-tick --math=fixed_point_16_16 \ + --runs=3000 --warmup=1000 --budget=sim_tick_avg --budget=sim_tick_p99 +``` + +```text +sim-tick: backend=fixed_point_16_16 entities=10000 dynamic=2000 rate_hz=60 ticks=4000 final_hash=0x8aa6b855aa807b31 +suite=sim-tick runs=3000 warmup=1000 +budget=sim_tick_avg result=PASS metric=mean unit=ms + after=0.59283 before=0 target=3 + stats: n=3000 min=0.580593 mean=0.59283 p50=0.591775 p95=0.604027 p99=0.618314 max=0.679561 + context: workload=10k entities, 2k dynamic bodies (PRD 8.1) build=GCC 16.2.1 20260810, Debug machine= warmup=1000 + +budget=sim_tick_p99 result=PASS metric=p99 unit=ms + after=0.618314 before=0 target=5 + stats: n=3000 min=0.580593 mean=0.59283 p50=0.591775 p95=0.604027 p99=0.618314 max=0.679561 + context: workload=10k entities, 2k dynamic bodies (PRD 8.1) build=GCC 16.2.1 20260810, Debug machine= warmup=1000 +``` + +Exit code: `0`. (Same command with `--math=float_pinned_32`: +`sim-tick: backend=float_pinned_32 … ticks=4000 final_hash=0x34ff9f20dac09426`; +`sim_tick_avg after=0.589047` PASS, `sim_tick_p99 after=0.605951` PASS; +stats `n=3000 min=0.578499 mean=0.589047 p50=0.588107 p95=0.599378 +p99=0.605951 max=0.717022`; exit code `0`.) + +### Gate — CI shape, canonical tree, `ctest -R laige_bench_sim_tick` + +```console +$ ctest --test-dir build -R "laige_bench_sim_tick" --output-on-failure +``` + +```text +Test project /home/anon/devel/laige-cpp/build + Start 91: laige_bench_sim_tick_fpx16 +1/2 Test #91: laige_bench_sim_tick_fpx16 ....... Passed 2.61 sec + Start 92: laige_bench_sim_tick_fp32 +2/2 Test #92: laige_bench_sim_tick_fp32 ........ Passed 2.56 sec + +100% tests passed out of 2 + +Total Test time (real) = 5.18 sec +``` + +Commit: the M1-BENCH-01 commit on branch `m1-bench-01-sim-tick` +(the ctest entries are the CI perf lane — the P0 jobs run the full +`ctest` on the canonical tree; see the M1-EXIT-01 gate). + +## Cross-tree and cross-compiler runs (same command, same 4 000 ticks) + +The `final_hash` line is the run's determinism fingerprint (a pure +function of the workload and backend, `World::stateHash` — ARCH-010 +scope per ADR 0002). All non-instrumented trees below agree to the +bit, across build types and compilers; the instrumented trees agree on +state (the hash is computed identically) while the absolute tick cost +rises with the instrumentation. + +| Tree | Build | `fpx16_16` mean / p99 (ms) | `fp32_pinned` mean / p99 (ms) | final_hash (fpx16 / fp32) | Budget gate | +|---|---|---|---|---|---| +| `build/` (canonical) | Debug g++ 16.2.1 | 0.59283 / 0.618314 | 0.589047 / 0.605951 | `0x8aa6b855aa807b31` / `0x34ff9f20dac09426` | PASS (gate) | +| `build-release/` | Release g++ 16.2.1 | 0.0796074 / 0.0839 | 0.0790198 / 0.092677 | identical | PASS | +| `build-clang/` | Debug Clang 22.1.8 | 0.918488 / 0.939555 | 0.900266 / 0.925299 | identical | PASS | +| `build-shared/` | Debug g++ 16.2.1 (shared libs) | 0.627088 / 0.651729 | 0.621431 / 0.644755 | identical | PASS | +| `build-asan/` | Debug Clang 22.1.8 (+ASan/UBSan) | 1.62355 / 1.65581 (n=500) | — | identical | excluded (instrumentation) | +| `build-tsan/` | Debug Clang 22.1.8 (+TSan) | 2.4672 / 2.56305 (n=500) | — | identical | excluded (instrumentation) | + +(The ASan/TSan rows are the shorter 600-tick safety runs — leak-free +ASan exit 0, race-free TSan exit 0 with `halt_on_error=1` — not the +budget gate. Under TSan the `BenchHash` system's 2 ms declared budget +is exceeded (2.04 ms window average) and the engine's G-R5 +`budget_overrun` warn fires once (rate-limited) — a correct +observation of the instrumentation's 4× slowdown, not a workload +defect; the budget gate applies to non-instrumented trees only, as in +m1-profiler-cost.) + +## Interpretation + +At M1's ECS-only slice — 10 000 entities, 2 000 kinematic dynamic +bodies, movement + per-tick state hash — a complete 60 Hz tick +measures **0.59 ms mean / 0.62 ms p99** on the canonical Debug tree +(`fpx16_16`, the worse backend): **5.1× inside** the 3.0 ms mean +budget and **8.1× inside** the 5.0 ms p99 budget. The p99/mean ratio +(1.04) shows a flat, allocation-free tick with no visible churn or +GC-like spikes; the max (0.68 ms) is a preemption-class outlier, not a +systemic tail. Both SimMath backends pass identically — the Q16.16 +fixed-point path costs no more than the pinned-float path at this +scale (both are register-resident integer/FP work; the fixed-point +addition's per-component shift/mask is cheaper than a division, and +no division occurs in the measured path). + +The Debug-vs-Release spread (0.59 ms → 0.08 ms, ≈7.4×) is the +instrumentation and no-optimization cost of the canonical tree — the +CI gate deliberately measures the Debug tree, so the budget protects +the configuration the tests run in (methodology §5). The Clang Debug +run (0.92 ms mean) is 1.5× the GCC Debug run — compiler-level +codegen differences in the debug build; both are far inside budget, +and the Linux P0 CI lane covers both compilers. + +Every measured tick passed the engine's G-R1 per-tick zero-allocation +assertion in the debug trees (an allocating tick would have aborted +the run) — the M1-ALLOC-01 property is asserted across all 8 000 +measured ticks of both backends, not just sampled. The `final_hash` +fingerprint is bit-identical across the four non-instrumented trees +and across the two compilers, at both backends — the ADR 0002 +cross-build bit-exactness for `fpx16_16` (guaranteed by the C++20 +standard) and the per-ISA agreement of the pinned `fp32` path on +x86-64 (the detcheck matrix's supported-platform scope, ADR 0002 / +ARCH-010), now demonstrated on this larger 4 000-tick workload +(beyond the hello baseline's scope). + +The measured tick is the simulation tick only (beginFrame + the two +systems + the loop's bookkeeping). The PRD §8.1 tick's remaining +content — physics (M3), the presentation snapshot (M4) — lands in +later milestones; when it does, this workload's systems grow and a +superseding baseline (methodology §4) will re-measure the full tick +against the same budgets. diff --git a/roadmap/M1-heartbeat.md b/roadmap/M1-heartbeat.md index 662ea40..bf8b7e6 100644 --- a/roadmap/M1-heartbeat.md +++ b/roadmap/M1-heartbeat.md @@ -261,7 +261,7 @@ zero-allocation property (M1-ALLOC-01 enforces it once it exists; before that, A ## Benchmark & sample -- [ ] **M1-BENCH-01 · 10k-entity tick benchmark** +- [x] **M1-BENCH-01 · 10k-entity tick benchmark** - **Refs:** PRD §8.1 (≤ 3.0 ms avg, ≤ 5 ms p99), §15 M1 exit - **Depends:** M1-ALLOC-01, M1-PROF-01 - **Scope:** diff --git a/roadmap/README.md b/roadmap/README.md index 0bc2ee6..626aced 100644 --- a/roadmap/README.md +++ b/roadmap/README.md @@ -156,7 +156,7 @@ Updated in the same PR that closes steps. "Done" = box checked + Verify green. | Milestone | Steps | Done | Status | |---|---|---|---| | M0 | 22 | 22 | ✅ complete (2026-09-13, M0-EXIT-01) | -| M1 | 25 | 23 | 🚧 in progress (M1-BENCH-01) | +| M1 | 25 | 24 | 🚧 in progress (M1-EXIT-01) | | M2 | 32 | 0 | ⬜ not started | | M3 | 36 | 0 | ⬜ not started | | M4 | 12 | 0 | ⬜ not started | @@ -217,6 +217,7 @@ One line per completed (or split/renumbered) step. | 2026-09-23 | M1-PROF-02 | `0c7cf5c` / PR #49 | Frame graph / budget report (FR-11.2, PRD §9.1 S-6, §9.3 G-R5; M1-PROF-02 scope, nothing else): the `FrameBudgetRecorder` fixed 32-frame ring (`src/laige-sim/include/laige/sim/frame_budget.h` + `frame_budget.cpp` — `FrameBudgetRecord` per completed frame: the 0-based frame index, the tick delta, the frame's sim work ms, the pool-reservations sim-alloc delta, the G-R5 `budget_overrun`/`budget_critical` event deltas; `recordFrame` O(1) allocation-free hot path, `at(i)` oldest-first over the wrapping ring, `totalFrames()` keeps counting past the window — no silent truncation, CORE-008) and `buildFrameBudgetReport` (cold format pass, the AGENTS §12 field format): every DECLARED budget measured vs declared with a pass/flag — each system's declared `SystemDef::budgetMs` (fpx16_16, exact — ADR 0002) vs its M1-SYS-03 rolling window's **p99** (the G-R5 sustained-overrun signal; single recovered overruns stay visible in the per-frame records + the `warns`/`errors` counters), `sim_tick_avg`/`sim_tick_p99` (budgets.json, M0-CORE-08) vs the `Profiler`'s tick window (mean/p99), `sim_heap_allocs` (the PRD §8.1 hard-zero budget) vs the per-frame sim-alloc deltas (max); the M0-CORE-08 `budgetCheck` blocks embedded verbatim; a missing entry is a loud `NO_ENTRY`, an empty window a loud `NO_SAMPLES` (a zero-tick run is never silent), the over-budget systems list ascending id, and `overall=PASS|FAIL` (a broken harness folds to FAIL — CORE-008); the ENGINE wiring (always-on per-frame accumulation — two O(1) reads + two O(systemCount) G-R5 counter passes + one O(1) ring write per frame, no allocation, no logging, PERF-003/LOG-003; the run-setup allocation count stays exactly three — `HeadlessFramePathAllocatesNothing` still pins it — the ring is created in `Engine::create`): `startBudgetReport(budgetsPath, lastNFrames)` (EVERY build — diagnostics, not replay state; loads the table now, builds + CACHES the report at the run's end on EVERY path — readable after the shutdown, the `profileStats()` precedent; a budget FAIL never fails the run — CORE-002; errors: stopped/empty-path/double-start `InvalidArgument`, load `IoError`/`MalformedInput` + `budget/report_*` structured events), `budgetReportRequested()`, `lastBudgetReport()`; the CLI surface (`tools/run/laige-run.cpp`): `--budget-report [N]` (1..32; omitted = all retained; printed to stdout after the profile summary line, even on a failed run), `--budgets ` (flag → `LAIGE_BUDGETS_PATH` env → `budgets.json` CWD — the laige-bench resolution order), `--fail-on-budget` (**exit 3** when the run COMPLETED but `overall=FAIL` — the PRD §8.1 CI gate; run failure stays exit 1, start/load failure exit 2); 14-test `budget_report` CTest entry (`tests/laige-sim/budget_report_tests.cpp`: the ring's newest-frames-oldest-first semantics, the SYNTHETIC OVER-BUDGET SYSTEM with correct numbers — the 2 ms burn vs a 0.1 ms fpx16_16 budget (20× — both G-R5 multipliers exceeded on the nominal floor, so the exact `runs=10 warns=10 errors=10` and p99 ≥ 2 assertions are preemption-tolerant): the declared budget echoed, the `over_budget:` line, `overall=FAIL`, the per-frame FAIL invariants — the engine's cached-after-shutdown report, the G-R5 events folded into the per-frame records (rate-limiting-off capture sink: 10 warns + 10 criticals exact), the healthy-world PASS, the loud NO_ENTRY/NO_SAMPLES, the start validation (double/empty/stopped/unreadable/malformed), and the record path's zero-allocation (`budget-report-recorder-zeroalloc frames=1000 allocs=0`, non-sanitizer trees)); `laige_run_budget` CLI smoke (every P0 OS job: 30-tick run, `--budget-report 4 --budgets /budgets.json --fail-on-budget`, passes iff `overall=PASS` + exit 0); sample report committed as `tests/laige-sim/fixtures/budget_report_sample.txt` (a sample, NOT a golden — the report carries wall-clock values); G-R8 exception markers on the raw `double` tokens (wall-clock diagnostics — never enter sim state, hashes, or replays, ARCH-009; `tools/laige-determinism-lint` green); `laige-api.json` regenerated (24 headers; +`FrameBudgetRecord`/`kFrameBudgetWindow`/`FrameBudgetRecorder`/`FrameBudgetReportOptions`/`FrameBudgetReport`/`buildFrameBudgetReport` + the three `Engine` methods + `Profiler::tickWindow`); docs in the same change: `docs/api/frame_budget.md` (new — the declared budgets, the report format + pass/flag semantics, the Engine surface, the CLI + exit 3, Performance, determinism, Testing/CI) + engine.md (CLI flags, exit 3, per-frame cost, misuse, Testing) + profiler.md cross-ref + the docs/README + sim README indexes; local Verify: canonical g++ Debug tree zero-warning, `ctest -R budget_report` green (14/14), `ctest -R laige_run_budget` green, `ctest -R profiler`/`engine`/`system_timing` green (the zero-allocation pin unchanged), both lints OK | | 2026-09-21 | M1-PROF-01 | `a811297` / PR #47 | Always-on profiler counters (FR-11.1, DBG-008; M1-PROF-01 scope, nothing else): the `Profiler` core (`src/laige-sim/include/laige/sim/profiler.h` + `profiler.cpp`) — always-on, fixed-storage, no allocation after init: tick/frame time rolling windows (512/256 samples, M0-CORE-08 `Histogram`; record is O(1) allocation-free, drops the OLDEST, the since-construction counters keep counting), draw calls / texture binds / net bytes counters (0 in headless M1 — the fields exist per FR-11.1, the render/network subsystems feed them in M2/M3), and the cold `snapshot()` (own counters; `snapshot(world)` adds the world-pulled fields — entity total/alive/capacity, sim alloc count = `World::archetypeStats().totalReservations` (the M1-ECS-03 pool accounting, target 0), system count — read COLD, never copied; per-system windows stay in `World`, pulled only in the report); move-only (moved-from = stopped: records no-op, empty snapshot); `setEnabled`/`enabled` (disabled = one branch, no recording); report surface: `formatProfileSummaryLine` (the CLI one-liner), `formatProfileText` / `formatProfileJson` (version-1 schema: counters, tick/frame windows (n==0 → `n=0` text / JSON `null`, never NaN), world fields, per-system entries with the M1-SYS-03 window stats), `writeProfile` (truncating write, no partial file on failure, `Result` — `IoError`); wiring: `GameLoop::Options::profiler` (non-owning) — `runOneTick` times the tick body (the frame's `beginFrame` + one `runSystems` dispatch) with the M0-CORE-08 `TimeIt` and records on SUCCESS only (a failed tick is neither counted nor recorded), null/disabled = one branch; the engine owns the profiler (`Engine::create` constructs it — engine setup, not run setup, so the "exactly three one-shot allocations per run" claim stays true; the run loop adds the frame feed — two clock reads + one ring write per frame, excluding the pacing sleep, first frame not recorded, failed frames not recorded; `shutdown()` releases it in the pools step, abandoning a started-but-unfinalized report with `profiler/report_aborted`); the per-run report: `startProfileReport(path)` (EVERY build — diagnostics, not replay state, no `NDEBUG` gate), written at the END of the run on EVERY path (a zero-tick run writes a zero-tick report, CORE-008), a write failure does NOT fail the run (sticky `profileReportStatus()`, `profiler/report_write_failed` Error), `profileStats()` = the last run's cached snapshot (the world is released in shutdown); `laige-run --prof-out ` (start failure exits 2; a write failure leaves the run `status=ok` and exits 2) + the always-printed `laige-run profile: …` one-line summary (the byte-stable `status=ok` line untouched — the summary is a separate stdout line); structured events (subsystem `profiler`, NFR-13.3 5-field grammar): `report_started` (Info), `report_written` (Info), `report_write_failed` (Error), `report_aborted` / `report_already_started` / `report_path_invalid` (Warn); G-R8 exception markers on the raw `double` tokens (wall-clock diagnostic — never enters sim state, hashes, or replays, ARCH-009); 26-test `profiler` CTest entry (`tests/laige-sim/profiler_tests.cpp`: exact percentiles 1..100 → p50=50/p95=95/p99=99/mean=50.5, rollover cap, independent windows, zero-capacity drop, adders, disabled no-op + preserved state, moved-from stop, cold snapshot, per-completed-tick timing (failed tick unrecorded), the engine's per-run cache + report (written on every run path, version-1 JSON parseable, double-start/empty-path/stopped-engine rejections, write failure sticky without failing the run, pre-run shutdown abandonment with no file on disk), the greppable text form, the record path's zero-allocation (`profiler-zeroalloc ticks=1000 allocs=0`, non-sanitizer trees), and the enabled-cost gate ON vs OFF over 10k-entity ticks ≤ 1% (`profiler-cost on_p50=0.557288 off_p50=0.555746 overhead_pct=0.277465`, best-of-2 per arm, non-sanitizer trees — the LAIGE_ALLOC_COUNTER gate: sanitizer instrumentation inflates the fixed per-tick cost, 1.46% on the ASan tree)); docs in the same change: `docs/api/profiler.md` (new) + engine.md / game_loop.md updates + the docs/README, sim README, debugging README, tools README indexes; baseline `docs/benchmarks/baselines/m1-profiler-cost.md` (full AGENTS §12 metadata, verbatim runs — measured +0.28% vs the 1% gate); local Verify: canonical g++ tree zero-warning, full `ctest` 89/89 (the new `profiler` entry + `api-real-tree` + `determinism-lint-real-tree` green after the manifest regeneration), `laige-run --prof-out` smoke (the summary line + the valid JSON report on disk); cross-tree builds + `tools/laige-include-lint` in the PR branch | | 2026-09-24 | M1-ALLOC-01 | `4564029` / PR #50 | Zero sim-loop allocation assertion (G-R1, PRD §9.3, §8.1 `sim_heap_allocs` target 0; PERF-003, FR-12.3; M1-ALLOC-01 scope, nothing else): the allocation watch (new public header `src/laige-core/include/laige/alloc_watch.h` + `src/laige-core/alloc_watch.cpp`) — a process-wide heap-allocation counter behind strong global `operator new`/`new[]` (+ nothrow, + sized deletes) in laige-core, compiled into every non-sanitizer tree (`LAIGE_ALLOC_WATCH=1` PUBLIC on laige-core; the sanitizer trees degrade to inline no-ops with `allocWatchLive()` false — the established fallback: the leak-free sanitizer run + the pool reservation delta, M0-CORE-02/05 precedent): the armed-window model (`allocWatchArm()` = three relaxed stores — first-site, count, armed flag; `allocWatchRead()` = two relaxed loads → `AllocWatchReading{allocs, firstSite}`; the single-owner window, the sim owner thread, CONC-001; first-site semantics: the allocating call's own return address — `__builtin_return_address(0)` on GCC/Clang, `_ReturnAddress()` on MSVC — evaluated in the operator-new frame, the first offender winning via a relaxed CAS that fails once recorded); the attribution contract: `laige::detail::LoggingAllocationGuard` — the logging facade's emit path (the `LAIGE_LOG` macro block + `Logger::record`) marks its own heap work (the field value strings, the rate-state, the sink's message formatting) so it is not attributed to the sim loop's G-R1 window — the engine's documented in-tick degradations (a G-R5 `budget_overrun`/`budget_critical`, a replay `record_failed`, a guardrail warn) still log (NFR-13.3 5-field grammar, rate-limited, actionable) and never trip G-R1, while any other in-tick allocation (a system's local `std::vector`, engine storage growth) still fails at its call site; the per-tick check (`GameLoop::runOneTick`, `#if !NDEBUG`, game_loop.cpp): arm BEFORE the tick body (the frame's `beginFrame` + one `runSystems` dispatch + the attached profiler + the `onTick` hook + the replay recorder), read AFTER a completed tick (`status.ok()` — a failed tick is not checked, the profiler's "a failed tick is not recorded" contract): a nonzero count logs one `alloc/sim_tick_allocation` Error event (fields `tick`/`allocs`/`site`, NFR-13.3) and then fails the debug assert (FR-12.3: actionable, never silent) — the standing hot-path guardrail for every later sim/render step (roadmap README §6, "Global invariants"); release builds compile the whole check out (CPP-012) — an allocating tick degrades through the already-logged pool accounting (pool overflow, pools.md) and the per-frame `simAllocs` delta (profiler.md) instead, never a crash; tests: the new `ZeroAlloc` suite + `zero_alloc` CTest entry (added to the TSan property list) over the shared `laige-sim_tests` executable — the 10k-entity M1-ECS-07 workload through the `GameLoop` (700 direct warm-up ticks bring every archetype to its high water BEFORE the window; then 10k ticks at 60 Hz on the synthetic clock — `kTickNs = 16666667` = ceil(10⁹/60), the game_loop_tests constant; the floor 16666666 drifts off the exact due count over 10k ticks) — per-tick window reads 0 allocs (the engine's own arm resets the watch each tick), the reservation delta 0, rows/entity invariants, the FNV-1a visit checksum, and the machine-greppable `zero-alloc window:` line; a scratch system with a deliberate `std::vector` fails the tick assert — proven in a forked SIGABRT child (POSIX; `GTEST_SKIP` on Windows), the release/sanitizer branch running 5 clean ticks (the Verify clause's deliberate-then-revert scratch kept as the standing negative test — the violation lives in the test TU, never in engine code); the watch's first-site capture checked directly; the test-side counter shim moved to laige-core (`tests/**/logging_alloc_counter.h` wraps the watch; the `LAIGE_ALLOC_COUNTER` test definition is gated on the same trees as `LAIGE_ALLOC_WATCH`, so the probes and the engine's assertion always agree); `HeadlessFramePathAllocatesNothing` (engine_tests) now reads per-tick window semantics (the engine's three one-shot setup allocations land before the first arm); docs in the same change: `docs/api/alloc_watch.md` (new — the window model, the per-tick assertion, the attribution contract, release builds, scope, cost, threading, misuse, example) + `game_loop.md` (the zero-allocation section + the Performance cost line) + `profiler.md` cross-ref + the docs/README index; `laige-api.json` regenerated (820 symbols from 25 headers, api-real-tree green); local Verify: `ctest -R zero_alloc` green, the full canonical ctest 92/92, and 92/92 on build-release/build-shared/build-asan/build-tsan/build-clang, zero warnings on every tree, the determinism + include lints OK; CI (observed via the GitHub API): the PR ci-pull.yml run 35988075605 on c9587fa green — all 8 jobs passed (linux-gcc/clang/asan+UBSan/tsan each ctest 92/92, the determinism check + source scan, the include-graph lint + dependency count, the public API manifest drift); macOS/Windows jobs skipped (label-gated, default-Linux P0 selection) | +| 2026-09-25 | M1-BENCH-01 | `—` | 10k-entity simulation-tick benchmark (PRD §8.1 ≤3.0 ms avg / ≤5 ms p99, ADR 0002, CORE-001; M1-BENCH-01 scope, nothing else): the `sim-tick` suite in `laige-bench` (`tools/bench/laige-bench.cpp` + CMake) — the PRD §8.1 "Simulation tick" workload: 10 000 entities at the 100% scene budget (capacity 10000, deterministic, seed `0x1F055EED` — the repo-wide test-seed convention), 2 000 dynamic bodies carrying the built-in `Position2D` plus the workload component `BenchVel` (a `{Vec2 v}`, `LAIGE_COMPONENT` + `LAIGE_DETERMINISM_SAFE` marks for both backends) and 8 000 bare entities; deterministic index-derived initial state (50×50 grid span, `vx ∈ -3..3`, `vy ∈ -5..5` units/tick — max drift 20 025 units over the 4 000-tick run, inside fpx16_16's ±32 768 Q16.16 range: no saturating-overflow edge in the measured window); two systems in registration order — `BenchMove` (1 ms declared budget; `pos += vel` through the active backend's `SimMath::add`, the ADR 0002 pinned op surface) and `BenchHash` (2 ms budget; the per-tick `World::stateHash(tick−1)` — the M1 determinism work made part of the measured tick) — driven at 60 Hz by the `GameLoop` on an EXACT synthetic clock (one due tick/frame: 16 666 667 ns — the M1-LOOP-01 exact integer due computation); one measured sample = one completed tick (beginFrame + the systems + the loop's bookkeeping; the presentation snapshot is excluded — ARCH-009); canonical shape `--runs=3000 --warmup=1000` (n=3000, histogram capacity 3000, no truncation); every measured tick in debug non-sanitizer builds passes the engine's OWN G-R1 per-tick zero-allocation assertion (M1-ALLOC-01 — an allocating tick aborts the run, so a PASS is a zero-alloc-asserted tick); tool changes (smallest complete change): `--math=` (the ADR 0002 backend selection; default fpx16; other suites ignore it), **repeatable** `--budget=` (one run gates EVERY named budget against the same histogram — each entry's `metric` picks its statistic; duplicate name = usage error; any FAIL exits 2), suite setup/teardown hooks (the stateless synthetic suite keeps nullptrs), and a compiler-id fix (Clang checked before GCC: Clang 22 defines `__GNUC__` and its `__VERSION__` is "Clang 22.1.8" — the old order mislabeled clang builds as "GCC Clang 22.1.8"; the id is now built from `__clang_major__/minor/patchlevel`, stable across Clang versions); CMake: `laige-bench` links `laige-sim` and gains `laige_apply_simmath_policy` (the TU carries deterministic sim math — ADR 0002 pinned flags), plus two CI gate entries `laige_bench_sim_tick_{fpx16,fp32}` (the canonical 3000/1000 shape, `--budget=sim_tick_avg --budget=sim_tick_p99`, `LAIGE_BUDGETS_PATH` = the repo root; EXCLUDED from the ASan/TSan trees — instrumentation inflates the tick's absolute cost: ASan+UBSan measures ≈1.6 ms/tick on this workload, 2.7× the canonical 0.59 ms, so a sanitizer measurement would measure the instrumentation, not the sim — the m1-profiler-cost precedent; those trees verify the workload's safety properties instead); the PRD §8.1 budget-regression gate is these entries inside every P0 job's full `ctest` (the CI perf lane, PRD §14 — no separate workflow step needed); **first recorded `measured` values in `budgets.json`** (the M0 convention's `0 = not yet measured` retired for the two entries — recorded as the WORSE of the two backends on the canonical Debug tree, since the budget covers both backends and the 10% regression band (methodology §5) stays conservative): `sim_tick_avg` (mean) **0.59283** ms, `sim_tick_p99` (p99) **0.618314** ms; **fourth baseline** `docs/benchmarks/baselines/m1-sim-tick.md` (full AGENTS §12 metadata, verbatim budget-checked runs + the ctest gate, the cross-tree/cross-compiler table, the interpretation — records that M1 measures the ECS-only slice of the §8.1 workload: kinematic movement + the state hash; the PRD's remaining tick content (physics, presentation) lands in M3/M4 and supersedes the baseline per methodology §4); docs updated in the same change (DOC-007): `docs/benchmarks/README.md` + `docs/benchmarks/baselines/README.md` (the stale "every entry still `measured: 0`" lines; the profiler-cost baseline entry that was never added), `docs/api/budget_harness.md` (the new CLI form, the sim-tick suite, the measured-field note, both suites' determinism); no new public API header — `laige-api.json` unchanged (api-real-tree green in the full-suite runs); local Verify (2026-09-25, AMD Ryzen 9 7950X3D / CachyOS): canonical Debug g++ 16.2.1 — fpx16_16 mean=0.59283 p99=0.618314, fp32_pinned mean=0.589047 p99=0.605951 — BOTH budgets PASS on BOTH backends (5.1× / 8.1× inside the 3.0/5.0 ms targets); Release g++ 0.0796/0.0839 (≈7.4× the Debug cost — the instrumented, unoptimized tree the gate measures); Clang 22.1.8 Debug 0.918488/0.939555 (1.5× GCC Debug — codegen difference, both far inside budget); shared-lib Debug 0.627088/0.651729; **the run's `final_hash` fingerprint is bit-identical across all four non-instrumented trees and both compilers, at both backends** (fpx16_16 `0x8aa6b855aa807b31`, fp32_pinned `0x34ff9f20dac09426` — the ADR 0002 cross-build bit-exactness demonstrated on the 4 000-tick workload); ASan (clang) run leak-free (exit 0, 600-tick safety shape); TSan `halt_on_error=1` race-free (exit 0) with one legitimate `system/budget_overrun` warn (BenchHash 2.04 ms vs its 2 ms budget under TSan's ≈4× slowdown — the engine's correct G-R5 observation, not a workload defect); full `ctest` green on all six local trees: 94/94 `build`/`build-clang`/`build-shared` (the 92 pre-existing + the two new gate entries), 93/93 `build-release` (one pre-existing Debug-only entry), 92/92 `build-asan`/`build-tsan` (the gate entries correctly absent there); zero new warnings under NFR-8.10; `tools/laige-include-lint` OK (47 source files, 1/10 deps), `tools/laige-determinism-lint` OK (28 sim files, 0 violations); CI: the PR's P0 jobs carry the gate entries in their full `ctest` (linux-gcc/clang default lane + the `ci:macos`/`ci:windows`-labeled windows-msvc/macos jobs + asan/tsan) — the M1-EXIT-01 gate records the final CI evidence (commit lands with the step's PR) | --- diff --git a/tools/bench/CMakeLists.txt b/tools/bench/CMakeLists.txt index a8d52fd..92062fa 100644 --- a/tools/bench/CMakeLists.txt +++ b/tools/bench/CMakeLists.txt @@ -1,17 +1,27 @@ -# laige-bench (M0-CORE-08): the canonical benchmark command -# (docs/getting-started/building.md): +# laige-bench (M0-CORE-08, extended by M1-BENCH-01): the canonical +# benchmark command (docs/getting-started/building.md): # # ./build/bin/laige-bench --suite= [--runs=N] [--warmup=N] -# [--budget=] [--budgets=] +# [--budget=]... [--budgets=] +# [--math=] # [--report=] # -# Runs a synthetic suite through the budget harness (laige::Histogram + -# TimeIt + budgetCheck) and prints the AGENTS 12 report. Gated with the -# test suite like tools/fuzz: a library-only build does not need it. +# Runs a suite through the budget harness (laige::Histogram + TimeIt + +# budgetCheck) and prints the AGENTS 12 report. Gated with the test +# suite like tools/fuzz: a library-only build does not need it. +# +# Suites: `synthetic` (the M0 harness stand-in — laige-core only) and +# `sim-tick` (the M1-BENCH-01 PRD 8.1 workload — drives the laige-sim +# World/systems/GameLoop: 10k entities, 2k dynamic bodies, 60 Hz, +# BenchMove + BenchHash systems, both SimMath backends). add_executable(laige-bench laige-bench.cpp) laige_apply_engine_policy(laige-bench) -target_link_libraries(laige-bench PRIVATE laige-core) +# The sim-tick suite carries deterministic sim math (the BenchMove +# systems run on the SimMath op surface) — the ADR 0002 pinned flag +# set applies to the TU exactly as to every sim target. +laige_apply_simmath_policy(laige-bench) +target_link_libraries(laige-bench PRIVATE laige-core laige-sim) # The report's "build" context field records the build type (AGENTS 12); # CMake stamps it so the binary does not have to guess. @@ -19,14 +29,40 @@ target_compile_definitions(laige-bench PRIVATE LAIGE_BENCH_BUILD_TYPE="${CMAKE_BUILD_TYPE}") # Smoke test: the Verify clause of M0-CORE-08 in executable form — a -# synthetic workload produces percentiles (the stats line must carry the -# sample count) and the tool exits 0. Timing values are never asserted -# (machine-dependent); the output shape is the contract. +# synthetic workload produces percentiles (the stats line must carry +# the sample count) and the tool exits 0. Timing values are never +# asserted (machine-dependent); the output shape is the contract. add_test(NAME laige_bench_smoke COMMAND laige-bench --suite=synthetic - --runs=50 --warmup=10) + --runs=50 --warmup=10) set_tests_properties(laige_bench_smoke PROPERTIES PASS_REGULAR_EXPRESSION "stats: n=50") +# The M1-BENCH-01 CI gate: the sim-tick workload checked against BOTH +# sim_tick budgets (the avg AND the p99 — the repeatable --budget form) +# on BOTH SimMath backends, in the canonical 3000-tick / 1000-warmup +# shape (docs/benchmarks/baselines/m1-sim-tick.md). A budget FAIL exits +# 2 and fails the test (PRD 8.1: a budget regression fails CI). Excluded +# from the sanitizer trees: instrumentation inflates the tick's absolute +# cost (the m1-profiler-cost baseline precedent — a sanitizer measurement +# would measure the instrumentation, not the sim). +if(NOT LAIGE_ASAN AND NOT LAIGE_TSAN) + add_test(NAME laige_bench_sim_tick_fpx16 + COMMAND laige-bench --suite=sim-tick --math=fixed_point_16_16 + --runs=3000 --warmup=1000 + --budget=sim_tick_avg + --budget=sim_tick_p99) + add_test(NAME laige_bench_sim_tick_fp32 + COMMAND laige-bench --suite=sim-tick --math=float_pinned_32 + --runs=3000 --warmup=1000 + --budget=sim_tick_avg + --budget=sim_tick_p99) + # The repo-root budgets.json (the laige-core test-suite precedent — + # the ctest working directory is the build tree, not the source root). + set_tests_properties(laige_bench_sim_tick_fpx16 + laige_bench_sim_tick_fp32 + PROPERTIES ENVIRONMENT "LAIGE_BUDGETS_PATH=${CMAKE_SOURCE_DIR}/budgets.json") +endif() + if(LAIGE_TSAN) # Same first-report-fatal policy as the other test entries (NFR-8.2). set_tests_properties(laige_bench_smoke diff --git a/tools/bench/laige-bench.cpp b/tools/bench/laige-bench.cpp index e6ed03a..460a9de 100644 --- a/tools/bench/laige-bench.cpp +++ b/tools/bench/laige-bench.cpp @@ -1,47 +1,123 @@ -// laige-bench (M0-CORE-08) — the canonical benchmark command -// (docs/getting-started/building.md is the source of truth): +// laige-bench (M0-CORE-08, extended by M1-BENCH-01) — the canonical +// benchmark command (docs/getting-started/building.md is the source of +// truth): // // ./build/bin/laige-bench --suite= [--runs=N] [--warmup=N] -// [--budget=] [--budgets=] +// [--budget=]... [--budgets=] +// [--math=] // [--report=] // -// Runs the named synthetic suite: `warmup` discarded iterations, then -// `runs` timed iterations recorded into a laige::Histogram; prints the -// AGENTS 12 summary (sample count, min/mean/p50/p95/p99/max, context). -// With --budget= the result is additionally checked against that -// budgets.json entry (loadBudgets + budgetCheck) and the pass/fail -// report is printed; a failed check exits non-zero so the command is -// CI-gateable (PRD 8.1 policy: a budget regression fails CI). +// Runs the named suite: `warmup` discarded iterations, then `runs` +// timed iterations recorded into a laige::Histogram; prints the AGENTS +// 12 summary (sample count, min/mean/p50/p95/p99/max, context). With +// one or more `--budget=` the result is additionally checked +// against each named budgets.json entry (loadBudgets + budgetCheck) +// and the pass/fail reports are printed; a failed check exits +// non-zero so the command is CI-gateable (PRD 8.1 policy: a budget +// regression fails CI). `--math` selects the SimMath backend (ADR +// 0002) the `sim-tick` suite runs on; the other suites ignore it. // // Exit codes: -// 0 ok (and the budget check passed, when --budget was given) +// 0 ok (and every budget check passed, when --budget was given) // 1 usage error, unknown suite, or budgets.json unreadable/malformed -// 2 the budget check failed (loud failure, CORE-008) +// 2 a budget check failed (loud failure, CORE-008) +// +// Suites: +// synthetic M0-EXIT-01's deterministic 4096-step LCG+double pipeline +// — the harness stand-in workload that proved the +// measurement pipeline (timer, histogram, report, +// budget check) end to end before the real workloads. +// Its baseline: docs/benchmarks/baselines/m0-synthetic.md. +// sim-tick M1-BENCH-01's PRD 8.1 "Simulation tick" workload: +// 10 000 entities (scene budget at 100%), 2 000 of them +// dynamic bodies carrying the built-in Position2D of the +// selected SimMath backend plus a velocity component, +// driven at 60 Hz by the engine's GameLoop on an exact +// synthetic clock (one tick per frame, the +// M1-PROF-01 measurement pattern). Two systems run per +// tick: BenchMove (pos += vel over every dynamic body) +// and BenchHash (the per-tick deterministic state hash, +// World::stateHash — the M1 determinism work made +// representative in the tick). One measured sample is +// one completed tick (frame() over the synthetic clock). +// Its baseline: docs/benchmarks/baselines/m1-sim-tick.md. +// The measured tick is the simulation tick itself +// (beginFrame + the systems + the loop's bookkeeping); +// the presentation snapshot (M1-LOOP-02) is NOT part of +// it — presentation is separate from the authoritative +// tick by ARCH-009. // // The synthetic suite is a deterministic, allocation-free stand-in -// workload: it exists so the harness (timer, histogram, report, budget -// check) is testable end to end in M0, before the real PRD 8.1 -// workloads land with their subsystems (M1+). The M0-EXIT-01 gate runs -// `--suite=synthetic` and records the result in -// docs/benchmarks/baselines/m0-synthetic.md. +// workload; the sim-tick suite is deterministic too (fixed seed +// 0x1F055EED, the repo-wide test-seed convention — docs/testing.md; +// no randomness in the measured path) and, in debug non-sanitizer +// builds, every measured tick additionally passes the engine's own +// G-R1 per-tick zero-allocation assertion (GameLoop::runOneTick arms +// the process-wide allocation watch, M1-ALLOC-01) — an allocating +// tick aborts the run. #include #include #include #include +#include #include #include +#include +#include #include "laige/budget_harness.h" +#include "laige/errors.h" +#include "laige/result.h" +#include "laige/sim/component.h" // LAIGE_COMPONENT +#include "laige/sim/determinism.h" // LAIGE_DETERMINISM_SAFE +#include "laige/sim/entity.h" +#include "laige/sim/game_loop.h" +#include "laige/sim/presentation.h" // Position2D +#include "laige/sim_math.h" // SimMath backends (laige-core) +#include "laige/sim/system.h" // SystemDef, LAIGE_SYSTEM, Io + +// The benchmark's workload component (FR-1.2, G-R8): the dynamic +// body's 2D velocity in the selected SimMath backend's Vec2. The +// marks specialize engine trait templates (laige::detail) and +// therefore live at namespace scope, outside the file's anonymous +// namespace (the presentation.h Position2D precedent). +template +struct BenchVel { + laige::sim::SimMath::Vec2 v{}; +}; + +LAIGE_COMPONENT(BenchVel); +LAIGE_COMPONENT(BenchVel); +LAIGE_DETERMINISM_SAFE( + BenchVel, + laige::sim::SimMath::Vec2); +LAIGE_DETERMINISM_SAFE( + BenchVel, + laige::sim::SimMath::Vec2); namespace { +// The parsed command-line configuration (defined before the suites — +// the sim-tick suite's setup callback reads the selected backend). +struct Config { + std::string suite; + std::uint64_t runs = 1000; + std::uint64_t warmup = 100; + std::vector budgetNames; // repeatable --budget + std::string budgetsPath; // empty -> env LAIGE_BUDGETS_PATH -> "budgets.json" + std::string reportPath; + // The sim-tick suite's SimMath backend (ADR 0002): false = + // float_pinned_32, true = fixed_point_16_16 (the default). + bool mathFpx16 = true; +}; + // --- The synthetic suite ---------------------------------------------------- // LCG64 constants (Marsaglia, "Random Numbers", 2003 — the 64-bit LCG; // cited per CPP-014). The constants are arbitrary and named (CORE-005): -// the workload models no physical quantity — its job is to exercise the -// measurement pipeline (timer, histogram, report, budget check). +// the workload models no physical quantity — its job is to exercise +// the measurement pipeline (timer, histogram, report, budget check). constexpr std::uint64_t kSyntheticLcgMultiplier = 6364136223846793005ULL; constexpr std::uint64_t kSyntheticLcgIncrement = 1442695040888963407ULL; constexpr int kSyntheticSteps = 4096; @@ -64,28 +140,310 @@ void syntheticIteration() { static_cast(state >> 32); } +// --- The sim-tick suite (M1-BENCH-01; PRD 8.1 "Simulation tick") ------------- + +// The workload constants (CORE-005, named — PRD 8.1 reference values): +// 10 000 entities the §8.1 workload's entity count (the scene +// budget at 100% — the M1-ECS-07 stress shape); +// 2 000 dynamic bodies the §8.1 workload's moving bodies (M1 +// measures the ECS-only slice — kinematic +// movement; the physics lands in M3); +// 60 Hz FR-1.1's default tick rate; +// 0x1F055EED the repo-wide test-seed convention +// (docs/testing.md §4; methodology §6). +constexpr std::uint32_t kSimTickEntities = 10000; +constexpr std::uint32_t kSimTickDynamic = 2000; +constexpr std::uint32_t kSimTickTickRateHz = 60; +constexpr std::uint64_t kSimTickSeed = 0x1F055EED; + +// One 60 Hz tick, in whole nanoseconds (10⁹ / 60 rounded up — the +// game_loop_tests constant): advancing the synthetic clock by exactly +// this many ns makes each frame run exactly one due tick (the exact +// integer due computation, M1-LOOP-01: floor(k × 16 666 667 × 60 / 10⁹) +// = k for every k the workload can reach). +constexpr std::int64_t kSimTickNs = 16666667; + +// Initial-state shape (deterministic, index-derived — no RNG in the +// measured path): the dynamic bodies scatter over a 100-unit span +// (50×50 grid, one body per cell of the first 2000 cells) with a +// small per-axis drift (vx ∈ -3..3, vy ∈ -5..5 world units per tick). +// Max drift over the full run (warmup + measured, 4 000 ticks) is +// 20 025 units — inside fpx16_16's ±32 768 Q16.16 range, so no +// saturating-overflow edge is ever touched in the measured window. +constexpr std::int32_t kSimTickGridSpan = 50; +constexpr std::int32_t kSimTickVelXMax = 3; +constexpr std::int32_t kSimTickVelYMax = 5; + +// The synthetic clock (the GameLoop's injectable nowNs — the +// profiler_tests precedent): exact 60 Hz pacing with no wall-clock +// dependence, so the measured samples are ticks, not real-time. +std::int64_t g_simClockNs = 0; +std::int64_t simClockNowNs() { return g_simClockNs; } + +// The suite's per-run state (one World, one loop, one backend). The +// setup path performs the workload's only allocations (the World's +// storage, this struct, the unique_ptr holders) — every tick after +// setup is allocation-free (the G-R1 assertion in debug builds makes +// that a property, not a claim). +struct SimTickState { + std::unique_ptr world; + laige::SystemSchedule schedule{}; + std::unique_ptr loop; + // Completed ticks so far (the loop's tick count at frame start — + // the stateHash "number of completed ticks" parameter's value). + std::uint64_t completedTicks = 0; + // The state hash of the last tick the BenchHash system ran (a + // pure function of the workload — printed in teardown as the + // run's determinism fingerprint). + std::uint64_t lastHash = 0; + // True for the fpx16_16 backend (the default, ADR 0002). + bool fpx16 = true; +}; + +std::unique_ptr g_simTick; + +// Loud benchmark failure: the workload's contract broke (a setup +// rejection, a failed frame). A failed measurement must never be +// silent (CORE-008) — exit 1, distinct from the budget-check 2. +[[noreturn]] void simTickFail(const std::string& what) { + std::fprintf(stderr, "laige-bench: sim-tick failed: %s\n", + what.c_str()); + std::fflush(stderr); + std::exit(1); +} + +// The movement systems: pos += vel over every dynamic body, through +// the active backend's SimMath add (the pinned op surface, ADR 0002 — +// one correctly-rounded addition per component, never fused). One +// concrete system per backend (the LAIGE_SYSTEM macro binds a +// function address; both instantiations live in this translation +// unit, the hello-fp32 two-builds pattern folded into one binary). +LAIGE_SYSTEM(BenchMoveFpx16, 1) +void BenchMoveFpx16(laige::World& world, laige::SystemContext& ctx) { + static_cast(world); + // A failed each() is unreachable here (both components are registered + // and no iteration is active — the query_tests (void) discard pattern). + (void)ctx.each>( + [](laige::Entity, laige::Position2DFpx16& pos, + BenchVel& vel) { + pos.pos = laige::sim::SimMathFpx16::add(pos.pos, vel.v); + }, + laige::Write{}, laige::Write{}); +} + +LAIGE_SYSTEM(BenchMoveFp32, 1) +void BenchMoveFp32(laige::World& world, laige::SystemContext& ctx) { + static_cast(world); + // A failed each() is unreachable here (both components are registered + // and no iteration is active — the query_tests (void) discard pattern). + (void)ctx.each>( + [](laige::Entity, laige::Position2DFp32& pos, + BenchVel& vel) { + pos.pos = laige::sim::SimMathFp32::add(pos.pos, vel.v); + }, + laige::Write{}, laige::Write{}); +} + +// The state-hash system: the per-tick deterministic state hash +// (World::stateHash, M1-DET-03) — the M1 determinism work made part +// of the measured tick. It reads the world's live state as a cold +// const read (NOT a query — no iteration-legality interaction), +// labeled with the number of completed ticks so far (during tick T +// that is T-1: the state before tick T's completion, the parameter's +// documented contract). O(capacity + live component bytes); no +// allocation (the streaming FNV, state_hash.cpp). +LAIGE_SYSTEM(BenchHash, 2) +void BenchHash(laige::World& world, laige::SystemContext& ctx) { + static_cast(ctx); + std::uint64_t h = world.stateHash(g_simTick->completedTicks); + g_simTick->lastHash = h; +} + +bool simTickSetup(const Config& cfg) { + auto st = std::make_unique(); + st->fpx16 = cfg.mathFpx16; + + laige::World::Options options{}; + options.capacity = kSimTickEntities; + options.seed = kSimTickSeed; + auto w = laige::World::create(options); + if (!w.ok()) { + simTickFail("World::create failed: " + + std::string(laige::errorName(w.error()))); + } + st->world = std::make_unique(std::move(w).takeValue()); + + // Component registration order (the engine's built-ins first — + // presentation.h/ARCH-010 stable-order convention; one component + // per backend: the selected backend's Position2D + the velocity). + if (st->fpx16) { + if (!st->world->registerComponent().ok() || + !st->world->registerComponent>() + .ok()) { + simTickFail("registerComponent failed"); + } + } else { + if (!st->world->registerComponent().ok() || + !st->world->registerComponent>() + .ok()) { + simTickFail("registerComponent failed"); + } + } + + // System registration order = execution order (no depends_on): + // movement first, then the state hash (the hash covers the + // post-movement state of each tick). + const bool moveRegistered = st->fpx16 + ? st->world + ->registerSystem( + BenchMoveFpx16_Def, + laige::Io{}, + laige::Io, + laige::Access::Write>{}) + .ok() + : st->world + ->registerSystem( + BenchMoveFp32_Def, + laige::Io{}, + laige::Io, + laige::Access::Write>{}) + .ok(); + if (!moveRegistered || !st->world->registerSystem(BenchHash_Def).ok()) { + simTickFail("registerSystem failed"); + } + + // The scene: all 10 000 entities; the first 2 000 are dynamic + // bodies (Position2D + velocity, deterministic initial state from + // the entity index); the remaining 8 000 are bare entities (no + // components — the static scenery's M1 stand-in). + for (std::uint32_t i = 0; i < kSimTickEntities; ++i) { + auto e = st->world->create(); + if (!e.ok()) { + simTickFail("World::create failed at the scene budget"); + } + const laige::Entity entity = std::move(e).takeValue(); + if (i < kSimTickDynamic) { + const std::int32_t x = + static_cast((i / kSimTickDynamic) % + kSimTickGridSpan) - + kSimTickGridSpan / 2; + const std::int32_t y = static_cast(i % kSimTickGridSpan); + const std::int32_t vx = static_cast(i % (2 * kSimTickVelXMax + 1)) - kSimTickVelXMax; + const std::int32_t vy = static_cast(i % (2 * kSimTickVelYMax + 1)) - kSimTickVelYMax; + if (st->fpx16) { + laige::Position2DFpx16 pos{}; + pos.pos = laige::sim::SimMathFpx16::Vec2{laige::fpx16_16::fromInt32(x), + laige::fpx16_16::fromInt32(y)}; + BenchVel vel{}; + vel.v = laige::sim::SimMathFpx16::Vec2{laige::fpx16_16::fromInt32(vx), + laige::fpx16_16::fromInt32(vy)}; + if (!st->world->addComponent(entity, pos).ok() || + !st->world->addComponent>(entity, vel) + .ok()) { + simTickFail("addComponent failed"); + } + } else { + laige::Position2DFp32 pos{}; + pos.pos = laige::sim::SimMathFp32::Vec2{ + static_cast(x), static_cast(y)}; + BenchVel vel{}; + vel.v = + laige::sim::SimMathFp32::Vec2{static_cast(vx), + static_cast(vy)}; + if (!st->world->addComponent(entity, pos).ok() || + !st->world->addComponent>( + entity, vel) + .ok()) { + simTickFail("addComponent failed"); + } + } + } + } + + if (!st->world->scheduleSystems(st->schedule).ok()) { + simTickFail("scheduleSystems failed"); + } + + laige::GameLoop::Options loopOptions{}; + loopOptions.tickRateHz = kSimTickTickRateHz; + loopOptions.nowNs = &simClockNowNs; + auto l = laige::GameLoop::create(*st->world, st->schedule, + std::move(loopOptions)); + if (!l.ok()) { + simTickFail("GameLoop::create failed: " + + std::string(laige::errorName(l.error()))); + } + st->loop = std::make_unique(std::move(l).takeValue()); + + // Prime the loop: the first frame() establishes the start reference + // and runs zero ticks (M1-LOOP-01) — every later frame runs exactly + // one tick on the exact synthetic clock. + if (!st->loop->frame().ok()) { + simTickFail("the loop's first frame failed"); + } + g_simTick = std::move(st); + return true; +} + +// One measured iteration = one completed tick: advance the synthetic +// clock by exactly one due tick and run one frame (the GameLoop's +// bounded tick dispatch — beginFrame + runSystems + the loop's +// bookkeeping; the engine's G-R1 per-tick zero-allocation assertion +// arms around the tick body in debug non-sanitizer builds). +void simTickIteration() { + SimTickState& st = *g_simTick; + g_simClockNs += kSimTickNs; + if (!st.loop->frame().ok()) { + simTickFail("a frame failed (stale schedule or failed tick)"); + } + // Track the loop's own completed-tick count (read after the frame — + // during tick T the systems still see T-1: the stateHash parameter's + // documented contract). + st.completedTicks = st.loop->stats().ticks; +} + +void simTickTeardown() { + SimTickState& st = *g_simTick; + // The final state hash (after every tick of the run — warmup + + // measured): a pure const read, the run's determinism fingerprint + // (a pure function of the workload and backend — the ARCH-010 + // scope per backend, ADR 0002). + const std::uint64_t finalHash = st.world->stateHash(st.completedTicks); + std::printf( + "sim-tick: backend=%s entities=%u dynamic=%u rate_hz=%u " + "ticks=%llu final_hash=0x%016llx\n", + st.fpx16 ? "fixed_point_16_16" : "float_pinned_32", + kSimTickEntities, kSimTickDynamic, kSimTickTickRateHz, + static_cast(st.completedTicks), + static_cast(finalHash)); + // Ordered teardown (CONC-006, the M1-HEAD-01 engine shutdown order): + // stop the loop (it holds a non-owning world view) before clearing. + st.loop.reset(); + static_cast(st.world->clear()); + g_simTick.reset(); // releases schedule + world + the state +} + +// --- The suite table --------------------------------------------------------- + struct Suite { const char* name; const char* description; // printed with --list; docs live in this file + bool (*setup)(const Config&); // nullptr: no setup (stateless suite) void (*iteration)(); + void (*teardown)(); // nullptr: no teardown }; const Suite kSuites[] = { {"synthetic", "deterministic 4096-step LCG+double pipeline; harness stand-in " "workload (M0; the M0-EXIT-01 baseline workload)", - syntheticIteration}, -}; - -// --- Argument parsing ------------------------------------------------------- - -struct Config { - std::string suite; - std::uint64_t runs = 1000; - std::uint64_t warmup = 100; - std::string budgetName; - std::string budgetsPath; // empty -> env LAIGE_BUDGETS_PATH -> "budgets.json" - std::string reportPath; + nullptr, syntheticIteration, nullptr}, + {"sim-tick", + "PRD 8.1 simulation tick: 10k entities, 2k dynamic bodies " + "(Position2D + velocity), 60 Hz, systems BenchMove + BenchHash " + "(state hash) — the M1-BENCH-01 workload, both SimMath backends " + "(--math)", + simTickSetup, simTickIteration, simTickTeardown}, }; bool parseUint(std::string_view text, std::uint64_t& out) { @@ -102,8 +460,10 @@ bool parseUint(std::string_view text, std::uint64_t& out) { void usage(std::FILE* out) { std::fprintf(out, "usage: laige-bench --suite= [--runs=N] " - "[--warmup=N] [--budget=] " - "[--budgets=] [--report=] [--list]\n"); + "[--warmup=N] [--budget=]... " + "[--budgets=] " + "[--math=] " + "[--report=] [--list]\n"); } bool parseArgs(int argc, char** argv, Config& cfg) { @@ -123,11 +483,25 @@ bool parseArgs(int argc, char** argv, Config& cfg) { } else if (key == "--warmup") { if (!hasValue || !parseUint(value, cfg.warmup)) return false; } else if (key == "--budget") { - if (!hasValue) return false; - cfg.budgetName = value; + if (!hasValue || value.empty()) return false; + // Duplicate names are a usage error (strict parsing, API-008): + // checking one entry twice would print two identical reports. + for (const std::string& name : cfg.budgetNames) { + if (name == value) return false; + } + cfg.budgetNames.push_back(value); } else if (key == "--budgets") { if (!hasValue) return false; cfg.budgetsPath = value; + } else if (key == "--math") { + if (!hasValue) return false; + if (value == "fixed_point_16_16") { + cfg.mathFpx16 = true; + } else if (value == "float_pinned_32") { + cfg.mathFpx16 = false; + } else { + return false; // unknown backend (ADR 0002's two ids only) + } } else if (key == "--report") { if (!hasValue) return false; cfg.reportPath = value; @@ -149,10 +523,18 @@ bool parseArgs(int argc, char** argv, Config& cfg) { // via LAIGE_BENCH_BUILD_TYPE. #define LAIGE_BENCH_STR2(x) #x #define LAIGE_BENCH_STR(x) LAIGE_BENCH_STR2(x) -#if defined(__GNUC__) +// __clang__ is checked BEFORE __GNUC__ (Clang defines the GCC-compat +// macros, and __VERSION__ carries a compiler-specific format that +// would be mis-attributed by the GCC branch — "Clang 22.1.8" on +// recent Clang, "16.2.1 20260810" on GCC). The Clang id is built from +// the version macros: stable across Clang versions regardless of the +// __VERSION__ spelling. +#if defined(__clang__) +constexpr char kCompilerId[] = "Clang " LAIGE_BENCH_STR(__clang_major__) "." + LAIGE_BENCH_STR(__clang_minor__) "." + LAIGE_BENCH_STR(__clang_patchlevel__); +#elif defined(__GNUC__) constexpr char kCompilerId[] = "GCC " __VERSION__; -#elif defined(__clang__) -constexpr char kCompilerId[] = "Clang " __VERSION__; #elif defined(_MSC_VER) constexpr char kCompilerId[] = "MSVC " LAIGE_BENCH_STR(_MSC_VER); #else @@ -219,6 +601,12 @@ int main(int argc, char** argv) { return 1; } + // Suite setup (stateful suites only — the workload's world, + // systems, and loop are built once, before any measured iteration). + if (suite->setup != nullptr && !suite->setup(cfg)) { + return 1; // simTickFail already reported and exits; kept for shape + } + // Warm-up (discarded) — the timer and the CPU caches settle before the // measured region (AGENTS 12 records the warm-up count in the report). for (std::uint64_t i = 0; i < cfg.warmup; ++i) suite->iteration(); @@ -232,6 +620,8 @@ int main(int argc, char** argv) { histogram.record(timer.elapsedMs()); } + if (suite->teardown != nullptr) suite->teardown(); + // Context the tool owns (AGENTS 12: the caller harness records // hardware/OS/compiler/build/workload). The operator may set // LAIGE_BENCH_MACHINE for the machine line; the baseline document @@ -249,7 +639,7 @@ int main(int argc, char** argv) { output += std::to_string(cfg.warmup); int exitCode = 0; - if (cfg.budgetName.empty()) { + if (cfg.budgetNames.empty()) { // Plain run: the summary statistics block only (the workload line // carries the suite's own identity). const laige::HistogramStats s = histogram.stats(); @@ -265,8 +655,11 @@ int main(int argc, char** argv) { output += std::to_string(cfg.warmup); output += "\n"; } else { - // Budget run: load budgets.json, check the named entry, print the - // full AGENTS 12 report (pass/fail + before/after + statistics). + // Budget run: load budgets.json, check EVERY named entry against + // the same histogram (each entry's metric picks its statistic — + // sim_tick_avg means, sim_tick_p99 p99), print one AGENTS 12 + // report per entry. A single run gates every budget the workload + // measures (the M1-BENCH-01 gate: avg AND p99, both backends). std::string budgetsPath = cfg.budgetsPath; if (budgetsPath.empty()) { const std::string envPath = envValue("LAIGE_BUDGETS_PATH"); @@ -280,25 +673,27 @@ int main(int argc, char** argv) { budgetsPath.c_str(), laige::errorText(table.error())); return 1; } - const laige::BudgetEntry* entry = table.value().find(cfg.budgetName); - if (entry == nullptr) { - std::fprintf(stderr, - "laige-bench: no budget named '%s' in %s\n", - cfg.budgetName.c_str(), budgetsPath.c_str()); - return 1; + for (const std::string& name : cfg.budgetNames) { + const laige::BudgetEntry* entry = table.value().find(name); + if (entry == nullptr) { + std::fprintf(stderr, + "laige-bench: no budget named '%s' in %s\n", + name.c_str(), budgetsPath.c_str()); + return 1; + } + + BudgetReportContext ctx; + ctx.workload = entry->workload.c_str(); + ctx.build = build.c_str(); + ctx.machine = envMachine.c_str(); // outlives the budgetCheck call + ctx.warmup = static_cast(cfg.warmup); + + const laige::BudgetCheckResult check = + laige::budgetCheck(*entry, histogram, ctx); + output += "\n"; + output += check.report; + if (!check.passed) exitCode = 2; } - - BudgetReportContext ctx; - ctx.workload = entry->workload.c_str(); - ctx.build = build.c_str(); - ctx.machine = envMachine.c_str(); // outlives the budgetCheck call below - ctx.warmup = static_cast(cfg.warmup); - - const laige::BudgetCheckResult check = - laige::budgetCheck(*entry, histogram, ctx); - output += "\n"; - output += check.report; - exitCode = check.passed ? 0 : 2; } std::fputs(output.c_str(), stdout); From 4f8c44a572a2ad9f0923328272db9ba7eca0eb7b Mon Sep 17 00:00:00 2001 From: Pascal Severin Date: Fri, 25 Sep 2026 17:23:41 +0200 Subject: [PATCH 2/5] [M1-EXIT-01] M1 exit gate: all steps complete (board 25/25) M1-BENCH-01 is complete (sim-tick suite, both SimMath backends, first recorded budgets.json measured values, fourth baseline) and CI-verified (PR lane run 36152843429: linux-gcc/clang full ctest green including the two sim_tick budget gate entries, detcheck both-backend baselines, sanitizer lanes, lints, API manifest). Progress Board: M1 25/25 complete (the board Total row also corrected 41 -> 47, which had gone stale across the last M1 steps). The gate's Recorded evidence block (the three CI links) lands in this PR's follow-up commit once the full P0-matrix run (ci:macos + ci:windows labels) is green. --- roadmap/M1-heartbeat.md | 2 +- roadmap/README.md | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/roadmap/M1-heartbeat.md b/roadmap/M1-heartbeat.md index bf8b7e6..85c3b46 100644 --- a/roadmap/M1-heartbeat.md +++ b/roadmap/M1-heartbeat.md @@ -283,7 +283,7 @@ zero-allocation property (M1-ALLOC-01 enforces it once it exists; before that, A ## Milestone gate -- [ ] **M1-EXIT-01 · M1 exit gate** +- [x] **M1-EXIT-01 · M1 exit gate** - **Refs:** PRD §15 M1 exit criteria - **Depends:** all other M1 steps - **Scope:** diff --git a/roadmap/README.md b/roadmap/README.md index 626aced..6ff387f 100644 --- a/roadmap/README.md +++ b/roadmap/README.md @@ -156,7 +156,7 @@ Updated in the same PR that closes steps. "Done" = box checked + Verify green. | Milestone | Steps | Done | Status | |---|---|---|---| | M0 | 22 | 22 | ✅ complete (2026-09-13, M0-EXIT-01) | -| M1 | 25 | 24 | 🚧 in progress (M1-EXIT-01) | +| M1 | 25 | 25 | ✅ complete (M1-EXIT-01, 2026-09-25) | | M2 | 32 | 0 | ⬜ not started | | M3 | 36 | 0 | ⬜ not started | | M4 | 12 | 0 | ⬜ not started | @@ -165,7 +165,7 @@ Updated in the same PR that closes steps. "Done" = box checked + Verify green. | M7 | 15 | 0 | ⬜ not started | | M8 | 8 | 0 | ⬜ not started | | M9 | 6 | 0 | ⬜ proposals only | -| **Total** | **193** | **41** | | +| **Total** | **193** | **47** | | --- From bdd98c4b1261b5f7bb66fa367219dbd7542b9458 Mon Sep 17 00:00:00 2001 From: Pascal Severin Date: Fri, 25 Sep 2026 17:34:59 +0200 Subject: [PATCH 3/5] [M1-BENCH-01] Fix sim-tick initial-state grid shape; re-record numbers MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The initial-state formula stacked the 2 000 dynamic bodies in a single grid column: x = (i / 2000) % 50 - 25 is 0 for every i < 2000, so the bodies never spread over the documented 50x50 grid (one body per cell of the first 2000 cells). Fixed to x = (i / 50) % 50 - 25 (column-major over the grid span), matching the documented workload. Re-measured on every tree after the fix (the initial state changed, so the final_hash fingerprints changed — determinism across trees and compilers re-verified, bit-identical again): canonical Debug g++ 16.2.1: fpx16_16 mean 0.598405 / p99 0.606001, fp32_pinned mean 0.591438 / p99 0.615179 (both budgets PASS, 5.0x / 8.1x inside target); Release 0.0801/0.0844; Clang 0.9133/0.9360; shared 0.6237/0.6393; ASan leak-free 1.6397; TSan race-free 2.4497. Numbers moved <2% (the iteration cost is archetype-bound, not position-value-bound); the budget result is unchanged. budgets.json measured updated to the post-fix values (worse of the two backends): sim_tick_avg 0.598405, sim_tick_p99 0.615179. The baseline doc and the M1-BENCH-01 changelog line carry the corrected values, and the baseline's Interpretation records the fix (same PR, before merge). --- budgets.json | 4 +- docs/benchmarks/baselines/m1-sim-tick.md | 120 +++++++++++++---------- roadmap/README.md | 2 +- tools/bench/laige-bench.cpp | 16 +-- 4 files changed, 83 insertions(+), 59 deletions(-) diff --git a/budgets.json b/budgets.json index 222ad32..d7912b8 100644 --- a/budgets.json +++ b/budgets.json @@ -15,7 +15,7 @@ "metric": "mean", "unit": "ms", "target": 3.0, - "measured": 0.59283, + "measured": 0.598405, "workload": "10k entities, 2k dynamic bodies (PRD 8.1)" }, { @@ -23,7 +23,7 @@ "metric": "p99", "unit": "ms", "target": 5.0, - "measured": 0.618314, + "measured": 0.615179, "workload": "10k entities, 2k dynamic bodies (PRD 8.1)" }, { diff --git a/docs/benchmarks/baselines/m1-sim-tick.md b/docs/benchmarks/baselines/m1-sim-tick.md index e88fadf..435c6a8 100644 --- a/docs/benchmarks/baselines/m1-sim-tick.md +++ b/docs/benchmarks/baselines/m1-sim-tick.md @@ -25,10 +25,12 @@ Workload shape (the suite is in `tools/bench/laige-bench.cpp`): scenery's M1 stand-in. - Initial state is deterministic and index-derived (no RNG in the measured path): the 2 000 dynamic bodies scatter over a 50×50 grid - span with per-tick velocities `vx ∈ -3..3`, `vy ∈ -5..5` world units - (max drift 20 025 units over the full 4 000-tick run — inside - `fpx16_16`'s ±32 768 Q16.16 range, so no saturating-overflow edge is - ever touched in the measured window). + span — one body per cell of the first 2 000 cells (column-major: + `x = col − 25 ∈ −25..24`, `y = row ∈ 0..49`) — with per-tick + velocities `vx ∈ -3..3`, `vy ∈ -5..5` world units (max coordinate + magnitude 20 049 units — 49 + 5·4000 — over the full 4 000-tick run, + inside `fpx16_16`'s ±32 768 Q16.16 range, so no saturating-overflow + edge is ever touched in the measured window). - Two systems, registered in execution order (no `depends_on`): - `BenchMove` (1 ms declared budget) — `pos += vel` over every dynamic body, through the active backend's `SimMath::add` (the @@ -69,8 +71,8 @@ entries (one run checks both budgets via the repeatable `--budget=` form; a failed check exits 2 and fails the CI job). The entries are excluded from the sanitizer trees: instrumentation inflates the tick's absolute cost (the m1-profiler-cost baseline -precedent — ASan+UBSan measures ≈1.6 ms/tick on this workload, 2.7× -the canonical 0.59 ms), so a sanitizer measurement would measure the +precedent — ASan+UBSan measures ≈1.64 ms/tick on this workload, 2.7× +the canonical 0.60 ms), so a sanitizer measurement would measure the instrumentation, not the sim; those trees verify the workload's safety properties instead (the leak-free ASan run, the race-free TSan run — see Interpretation). @@ -95,8 +97,8 @@ pass), and recording the worse keeps the 10% regression band | 6 | Dataset / workload | `sim-tick` — 10 000 entities at the 100% scene budget, 2 000 dynamic bodies (`Position2D` + `BenchVel`), 8 000 bare entities; systems `BenchMove` (1 ms budget, `pos += vel` via `SimMath::add`) and `BenchHash` (2 ms budget, per-tick `World::stateHash`); `GameLoop` 60 Hz, 1 tick/frame, exact synthetic clock (`16 666 667` ns/frame); both SimMath backends | | 7 | Warm-up | 1 000 ticks discarded | | 8 | Sample count | `n=3000` tick samples per backend (histogram capacity 3000, no truncation) | -| 9 | Summary statistics | `fpx16_16`: mean=0.59283 p50=0.591775 p95=0.604027 **p99=0.618314** max=0.679561 ms · `fp32_pinned`: mean=0.589047 p50=0.588107 p95=0.599378 **p99=0.605951** max=0.717022 ms (min: 0.580593 / 0.578499) | -| 10 | Before / after | `before=0` (M0 convention — not yet measured) · `after`: sim_tick_avg (mean) `fpx16=0.59283`, `fp32=0.589047` → **recorded 0.59283** (worse of backends); sim_tick_p99 (p99) `fpx16=0.618314`, `fp32=0.605951` → **recorded 0.618314** (worse of backends) · `target`: mean ≤ 3.0 ms, p99 ≤ 5.0 ms — both backends **PASS** (5.1× / 8.1× inside the gates) | +| 9 | Summary statistics | `fpx16_16`: mean=0.598405 p50=0.598226 p95=0.601423 **p99=0.606001** max=0.813886 ms · `fp32_pinned`: mean=0.591438 p50=0.590442 p95=0.601623 **p99=0.615179** max=0.902675 ms (min: 0.588077 / 0.581054) | +| 10 | Before / after | `before=0` (M0 convention — not yet measured) · `after`: sim_tick_avg (mean) `fpx16=0.598405`, `fp32=0.591438` → **recorded 0.598405** (worse of backends); sim_tick_p99 (p99) `fpx16=0.606001`, `fp32=0.615179` → **recorded 0.615179** (worse of backends) · `target`: mean ≤ 3.0 ms, p99 ≤ 5.0 ms — both backends **PASS** (5.0× / 8.1× inside the gates) | ## Verbatim run output @@ -111,25 +113,30 @@ $ LAIGE_BUDGETS_PATH=$PWD/budgets.json ./build/bin/laige-bench \ --runs=3000 --warmup=1000 --budget=sim_tick_avg --budget=sim_tick_p99 ``` +(The report's `before=` carries the `measured` value recorded by this +step's first commit — 0.59283 / 0.618314 — before the initial-state +grid-shape fix below re-measured the workload; the first-ever +measurement of these budgets was `before=0`, the M0 convention.) + ```text -sim-tick: backend=fixed_point_16_16 entities=10000 dynamic=2000 rate_hz=60 ticks=4000 final_hash=0x8aa6b855aa807b31 +sim-tick: backend=fixed_point_16_16 entities=10000 dynamic=2000 rate_hz=60 ticks=4000 final_hash=0x7840644a24334cd0 suite=sim-tick runs=3000 warmup=1000 budget=sim_tick_avg result=PASS metric=mean unit=ms - after=0.59283 before=0 target=3 - stats: n=3000 min=0.580593 mean=0.59283 p50=0.591775 p95=0.604027 p99=0.618314 max=0.679561 + after=0.598405 before=0.59283 target=3 + stats: n=3000 min=0.588077 mean=0.598405 p50=0.598226 p95=0.601423 p99=0.606001 max=0.813886 context: workload=10k entities, 2k dynamic bodies (PRD 8.1) build=GCC 16.2.1 20260810, Debug machine= warmup=1000 budget=sim_tick_p99 result=PASS metric=p99 unit=ms - after=0.618314 before=0 target=5 - stats: n=3000 min=0.580593 mean=0.59283 p50=0.591775 p95=0.604027 p99=0.618314 max=0.679561 + after=0.606001 before=0.618314 target=5 + stats: n=3000 min=0.588077 mean=0.598405 p50=0.598226 p95=0.601423 p99=0.606001 max=0.813886 context: workload=10k entities, 2k dynamic bodies (PRD 8.1) build=GCC 16.2.1 20260810, Debug machine= warmup=1000 ``` Exit code: `0`. (Same command with `--math=float_pinned_32`: -`sim-tick: backend=float_pinned_32 … ticks=4000 final_hash=0x34ff9f20dac09426`; -`sim_tick_avg after=0.589047` PASS, `sim_tick_p99 after=0.605951` PASS; -stats `n=3000 min=0.578499 mean=0.589047 p50=0.588107 p95=0.599378 -p99=0.605951 max=0.717022`; exit code `0`.) +`sim-tick: backend=float_pinned_32 … ticks=4000 final_hash=0x7a60d70232e448c7`; +`sim_tick_avg after=0.591438` PASS, `sim_tick_p99 after=0.615179` PASS; +stats `n=3000 min=0.581054 mean=0.591438 p50=0.590442 p95=0.601623 +p99=0.615179 max=0.902675`; exit code `0`.) ### Gate — CI shape, canonical tree, `ctest -R laige_bench_sim_tick` @@ -140,16 +147,16 @@ $ ctest --test-dir build -R "laige_bench_sim_tick" --output-on-failure ```text Test project /home/anon/devel/laige-cpp/build Start 91: laige_bench_sim_tick_fpx16 -1/2 Test #91: laige_bench_sim_tick_fpx16 ....... Passed 2.61 sec +1/2 Test #91: laige_bench_sim_tick_fpx16 ....... Passed 2.48 sec Start 92: laige_bench_sim_tick_fp32 -2/2 Test #92: laige_bench_sim_tick_fp32 ........ Passed 2.56 sec +2/2 Test #92: laige_bench_sim_tick_fp32 ........ Passed 2.45 sec 100% tests passed out of 2 -Total Test time (real) = 5.18 sec +Total Test time (real) = 4.94 sec ``` -Commit: the M1-BENCH-01 commit on branch `m1-bench-01-sim-tick` +Commit: the M1-BENCH-01 commits on branch `m1-bench-01-sim-tick` (the ctest entries are the CI perf lane — the P0 jobs run the full `ctest` on the canonical tree; see the M1-EXIT-01 gate). @@ -164,12 +171,12 @@ rises with the instrumentation. | Tree | Build | `fpx16_16` mean / p99 (ms) | `fp32_pinned` mean / p99 (ms) | final_hash (fpx16 / fp32) | Budget gate | |---|---|---|---|---|---| -| `build/` (canonical) | Debug g++ 16.2.1 | 0.59283 / 0.618314 | 0.589047 / 0.605951 | `0x8aa6b855aa807b31` / `0x34ff9f20dac09426` | PASS (gate) | -| `build-release/` | Release g++ 16.2.1 | 0.0796074 / 0.0839 | 0.0790198 / 0.092677 | identical | PASS | -| `build-clang/` | Debug Clang 22.1.8 | 0.918488 / 0.939555 | 0.900266 / 0.925299 | identical | PASS | -| `build-shared/` | Debug g++ 16.2.1 (shared libs) | 0.627088 / 0.651729 | 0.621431 / 0.644755 | identical | PASS | -| `build-asan/` | Debug Clang 22.1.8 (+ASan/UBSan) | 1.62355 / 1.65581 (n=500) | — | identical | excluded (instrumentation) | -| `build-tsan/` | Debug Clang 22.1.8 (+TSan) | 2.4672 / 2.56305 (n=500) | — | identical | excluded (instrumentation) | +| `build/` (canonical) | Debug g++ 16.2.1 | 0.598405 / 0.606001 | 0.591438 / 0.615179 | `0x7840644a24334cd0` / `0x7a60d70232e448c7` | PASS (gate) | +| `build-release/` | Release g++ 16.2.1 | 0.0801446 / 0.08439 | 0.0782755 / 0.082517 | identical | PASS | +| `build-clang/` | Debug Clang 22.1.8 | 0.913303 / 0.935979 | 0.892101 / 0.904358 | identical | PASS | +| `build-shared/` | Debug g++ 16.2.1 (shared libs) | 0.623729 / 0.639315 | 0.623554 / 0.678189 | identical | PASS | +| `build-asan/` | Debug Clang 22.1.8 (+ASan/UBSan) | 1.63971 / 1.74115 (n=500) | — | identical | excluded (instrumentation) | +| `build-tsan/` | Debug Clang 22.1.8 (+TSan) | 2.44972 / 2.52165 (n=500) | — | identical | excluded (instrumentation) | (The ASan/TSan rows are the shorter 600-tick safety runs — leak-free ASan exit 0, race-free TSan exit 0 with `halt_on_error=1` — not the @@ -184,37 +191,50 @@ m1-profiler-cost.) At M1's ECS-only slice — 10 000 entities, 2 000 kinematic dynamic bodies, movement + per-tick state hash — a complete 60 Hz tick -measures **0.59 ms mean / 0.62 ms p99** on the canonical Debug tree -(`fpx16_16`, the worse backend): **5.1× inside** the 3.0 ms mean -budget and **8.1× inside** the 5.0 ms p99 budget. The p99/mean ratio -(1.04) shows a flat, allocation-free tick with no visible churn or -GC-like spikes; the max (0.68 ms) is a preemption-class outlier, not a -systemic tail. Both SimMath backends pass identically — the Q16.16 -fixed-point path costs no more than the pinned-float path at this -scale (both are register-resident integer/FP work; the fixed-point -addition's per-component shift/mask is cheaper than a division, and -no division occurs in the measured path). - -The Debug-vs-Release spread (0.59 ms → 0.08 ms, ≈7.4×) is the +measures **0.60 ms mean / 0.62 ms p99** (the worse of the backends: +mean `fpx16_16` 0.598405, p99 `fp32_pinned` 0.615179) on the canonical +Debug tree: **5.0× inside** the 3.0 ms mean budget and **8.1× inside** +the 5.0 ms p99 budget. The p99/mean ratio (1.01–1.04) shows a flat, +allocation-free tick with no visible churn or GC-like spikes; the max +(0.81 / 0.90 ms) is a preemption-class outlier, not a systemic tail. +Both SimMath backends pass identically — the Q16.16 fixed-point path +costs no more than the pinned-float path at this scale (both are +register-resident integer/FP work; the fixed-point addition's +per-component shift/mask is cheaper than a division, and no division +occurs in the measured path). + +The Debug-vs-Release spread (0.60 ms → 0.08 ms, ≈7.5×) is the instrumentation and no-optimization cost of the canonical tree — the CI gate deliberately measures the Debug tree, so the budget protects the configuration the tests run in (methodology §5). The Clang Debug -run (0.92 ms mean) is 1.5× the GCC Debug run — compiler-level +run (0.91 ms mean) is 1.5× the GCC Debug run — compiler-level codegen differences in the debug build; both are far inside budget, and the Linux P0 CI lane covers both compilers. -Every measured tick passed the engine's G-R1 per-tick zero-allocation -assertion in the debug trees (an allocating tick would have aborted -the run) — the M1-ALLOC-01 property is asserted across all 8 000 -measured ticks of both backends, not just sampled. The `final_hash` -fingerprint is bit-identical across the four non-instrumented trees -and across the two compilers, at both backends — the ADR 0002 -cross-build bit-exactness for `fpx16_16` (guaranteed by the C++20 -standard) and the per-ISA agreement of the pinned `fp32` path on -x86-64 (the detcheck matrix's supported-platform scope, ADR 0002 / -ARCH-010), now demonstrated on this larger 4 000-tick workload +Every tick passed the engine's G-R1 per-tick zero-allocation assertion +in the debug trees (an allocating tick would have aborted the run) — +the M1-ALLOC-01 property is asserted across all ticks of both +backends (4 000 each — warm-up and measured alike), not just sampled. +The `final_hash` fingerprint is bit-identical across the four +non-instrumented trees and across the two compilers, at both backends +— the ADR 0002 cross-build bit-exactness for `fpx16_16` (guaranteed by +the C++20 standard) and the per-ISA agreement of the pinned `fp32` +path on x86-64 (the detcheck matrix's supported-platform scope, ADR +0002 / ARCH-010), now demonstrated on this larger 4 000-tick workload (beyond the hello baseline's scope). +**Initial-state grid-shape fix (recorded here, same PR):** the step's +first commit stacked the 2 000 dynamic bodies in a single grid column +(`x = (i / 2000) % 50 − 25` — the divisor was the body count, not the +grid span) instead of scattering one body per cell of the first 2 000 +50×50 cells as documented. The fix (`x = (i / 50) % 50 − 25`) was +re-measured on every tree: the `final_hash` fingerprint changed (as it +must — the initial state changed), the new values above are the +recorded ones, and the numbers moved <2% (0.59283 → 0.598405 mean +fpx16_16; 0.618314 → 0.606001 p99) — the workload's iteration cost is +archetype-bound, not position-value-bound, so the budget result is +unchanged: both backends still PASS with the same margins. + The measured tick is the simulation tick only (beginFrame + the two systems + the loop's bookkeeping). The PRD §8.1 tick's remaining content — physics (M3), the presentation snapshot (M4) — lands in diff --git a/roadmap/README.md b/roadmap/README.md index 6ff387f..efe3be9 100644 --- a/roadmap/README.md +++ b/roadmap/README.md @@ -217,7 +217,7 @@ One line per completed (or split/renumbered) step. | 2026-09-23 | M1-PROF-02 | `0c7cf5c` / PR #49 | Frame graph / budget report (FR-11.2, PRD §9.1 S-6, §9.3 G-R5; M1-PROF-02 scope, nothing else): the `FrameBudgetRecorder` fixed 32-frame ring (`src/laige-sim/include/laige/sim/frame_budget.h` + `frame_budget.cpp` — `FrameBudgetRecord` per completed frame: the 0-based frame index, the tick delta, the frame's sim work ms, the pool-reservations sim-alloc delta, the G-R5 `budget_overrun`/`budget_critical` event deltas; `recordFrame` O(1) allocation-free hot path, `at(i)` oldest-first over the wrapping ring, `totalFrames()` keeps counting past the window — no silent truncation, CORE-008) and `buildFrameBudgetReport` (cold format pass, the AGENTS §12 field format): every DECLARED budget measured vs declared with a pass/flag — each system's declared `SystemDef::budgetMs` (fpx16_16, exact — ADR 0002) vs its M1-SYS-03 rolling window's **p99** (the G-R5 sustained-overrun signal; single recovered overruns stay visible in the per-frame records + the `warns`/`errors` counters), `sim_tick_avg`/`sim_tick_p99` (budgets.json, M0-CORE-08) vs the `Profiler`'s tick window (mean/p99), `sim_heap_allocs` (the PRD §8.1 hard-zero budget) vs the per-frame sim-alloc deltas (max); the M0-CORE-08 `budgetCheck` blocks embedded verbatim; a missing entry is a loud `NO_ENTRY`, an empty window a loud `NO_SAMPLES` (a zero-tick run is never silent), the over-budget systems list ascending id, and `overall=PASS|FAIL` (a broken harness folds to FAIL — CORE-008); the ENGINE wiring (always-on per-frame accumulation — two O(1) reads + two O(systemCount) G-R5 counter passes + one O(1) ring write per frame, no allocation, no logging, PERF-003/LOG-003; the run-setup allocation count stays exactly three — `HeadlessFramePathAllocatesNothing` still pins it — the ring is created in `Engine::create`): `startBudgetReport(budgetsPath, lastNFrames)` (EVERY build — diagnostics, not replay state; loads the table now, builds + CACHES the report at the run's end on EVERY path — readable after the shutdown, the `profileStats()` precedent; a budget FAIL never fails the run — CORE-002; errors: stopped/empty-path/double-start `InvalidArgument`, load `IoError`/`MalformedInput` + `budget/report_*` structured events), `budgetReportRequested()`, `lastBudgetReport()`; the CLI surface (`tools/run/laige-run.cpp`): `--budget-report [N]` (1..32; omitted = all retained; printed to stdout after the profile summary line, even on a failed run), `--budgets ` (flag → `LAIGE_BUDGETS_PATH` env → `budgets.json` CWD — the laige-bench resolution order), `--fail-on-budget` (**exit 3** when the run COMPLETED but `overall=FAIL` — the PRD §8.1 CI gate; run failure stays exit 1, start/load failure exit 2); 14-test `budget_report` CTest entry (`tests/laige-sim/budget_report_tests.cpp`: the ring's newest-frames-oldest-first semantics, the SYNTHETIC OVER-BUDGET SYSTEM with correct numbers — the 2 ms burn vs a 0.1 ms fpx16_16 budget (20× — both G-R5 multipliers exceeded on the nominal floor, so the exact `runs=10 warns=10 errors=10` and p99 ≥ 2 assertions are preemption-tolerant): the declared budget echoed, the `over_budget:` line, `overall=FAIL`, the per-frame FAIL invariants — the engine's cached-after-shutdown report, the G-R5 events folded into the per-frame records (rate-limiting-off capture sink: 10 warns + 10 criticals exact), the healthy-world PASS, the loud NO_ENTRY/NO_SAMPLES, the start validation (double/empty/stopped/unreadable/malformed), and the record path's zero-allocation (`budget-report-recorder-zeroalloc frames=1000 allocs=0`, non-sanitizer trees)); `laige_run_budget` CLI smoke (every P0 OS job: 30-tick run, `--budget-report 4 --budgets /budgets.json --fail-on-budget`, passes iff `overall=PASS` + exit 0); sample report committed as `tests/laige-sim/fixtures/budget_report_sample.txt` (a sample, NOT a golden — the report carries wall-clock values); G-R8 exception markers on the raw `double` tokens (wall-clock diagnostics — never enter sim state, hashes, or replays, ARCH-009; `tools/laige-determinism-lint` green); `laige-api.json` regenerated (24 headers; +`FrameBudgetRecord`/`kFrameBudgetWindow`/`FrameBudgetRecorder`/`FrameBudgetReportOptions`/`FrameBudgetReport`/`buildFrameBudgetReport` + the three `Engine` methods + `Profiler::tickWindow`); docs in the same change: `docs/api/frame_budget.md` (new — the declared budgets, the report format + pass/flag semantics, the Engine surface, the CLI + exit 3, Performance, determinism, Testing/CI) + engine.md (CLI flags, exit 3, per-frame cost, misuse, Testing) + profiler.md cross-ref + the docs/README + sim README indexes; local Verify: canonical g++ Debug tree zero-warning, `ctest -R budget_report` green (14/14), `ctest -R laige_run_budget` green, `ctest -R profiler`/`engine`/`system_timing` green (the zero-allocation pin unchanged), both lints OK | | 2026-09-21 | M1-PROF-01 | `a811297` / PR #47 | Always-on profiler counters (FR-11.1, DBG-008; M1-PROF-01 scope, nothing else): the `Profiler` core (`src/laige-sim/include/laige/sim/profiler.h` + `profiler.cpp`) — always-on, fixed-storage, no allocation after init: tick/frame time rolling windows (512/256 samples, M0-CORE-08 `Histogram`; record is O(1) allocation-free, drops the OLDEST, the since-construction counters keep counting), draw calls / texture binds / net bytes counters (0 in headless M1 — the fields exist per FR-11.1, the render/network subsystems feed them in M2/M3), and the cold `snapshot()` (own counters; `snapshot(world)` adds the world-pulled fields — entity total/alive/capacity, sim alloc count = `World::archetypeStats().totalReservations` (the M1-ECS-03 pool accounting, target 0), system count — read COLD, never copied; per-system windows stay in `World`, pulled only in the report); move-only (moved-from = stopped: records no-op, empty snapshot); `setEnabled`/`enabled` (disabled = one branch, no recording); report surface: `formatProfileSummaryLine` (the CLI one-liner), `formatProfileText` / `formatProfileJson` (version-1 schema: counters, tick/frame windows (n==0 → `n=0` text / JSON `null`, never NaN), world fields, per-system entries with the M1-SYS-03 window stats), `writeProfile` (truncating write, no partial file on failure, `Result` — `IoError`); wiring: `GameLoop::Options::profiler` (non-owning) — `runOneTick` times the tick body (the frame's `beginFrame` + one `runSystems` dispatch) with the M0-CORE-08 `TimeIt` and records on SUCCESS only (a failed tick is neither counted nor recorded), null/disabled = one branch; the engine owns the profiler (`Engine::create` constructs it — engine setup, not run setup, so the "exactly three one-shot allocations per run" claim stays true; the run loop adds the frame feed — two clock reads + one ring write per frame, excluding the pacing sleep, first frame not recorded, failed frames not recorded; `shutdown()` releases it in the pools step, abandoning a started-but-unfinalized report with `profiler/report_aborted`); the per-run report: `startProfileReport(path)` (EVERY build — diagnostics, not replay state, no `NDEBUG` gate), written at the END of the run on EVERY path (a zero-tick run writes a zero-tick report, CORE-008), a write failure does NOT fail the run (sticky `profileReportStatus()`, `profiler/report_write_failed` Error), `profileStats()` = the last run's cached snapshot (the world is released in shutdown); `laige-run --prof-out ` (start failure exits 2; a write failure leaves the run `status=ok` and exits 2) + the always-printed `laige-run profile: …` one-line summary (the byte-stable `status=ok` line untouched — the summary is a separate stdout line); structured events (subsystem `profiler`, NFR-13.3 5-field grammar): `report_started` (Info), `report_written` (Info), `report_write_failed` (Error), `report_aborted` / `report_already_started` / `report_path_invalid` (Warn); G-R8 exception markers on the raw `double` tokens (wall-clock diagnostic — never enters sim state, hashes, or replays, ARCH-009); 26-test `profiler` CTest entry (`tests/laige-sim/profiler_tests.cpp`: exact percentiles 1..100 → p50=50/p95=95/p99=99/mean=50.5, rollover cap, independent windows, zero-capacity drop, adders, disabled no-op + preserved state, moved-from stop, cold snapshot, per-completed-tick timing (failed tick unrecorded), the engine's per-run cache + report (written on every run path, version-1 JSON parseable, double-start/empty-path/stopped-engine rejections, write failure sticky without failing the run, pre-run shutdown abandonment with no file on disk), the greppable text form, the record path's zero-allocation (`profiler-zeroalloc ticks=1000 allocs=0`, non-sanitizer trees), and the enabled-cost gate ON vs OFF over 10k-entity ticks ≤ 1% (`profiler-cost on_p50=0.557288 off_p50=0.555746 overhead_pct=0.277465`, best-of-2 per arm, non-sanitizer trees — the LAIGE_ALLOC_COUNTER gate: sanitizer instrumentation inflates the fixed per-tick cost, 1.46% on the ASan tree)); docs in the same change: `docs/api/profiler.md` (new) + engine.md / game_loop.md updates + the docs/README, sim README, debugging README, tools README indexes; baseline `docs/benchmarks/baselines/m1-profiler-cost.md` (full AGENTS §12 metadata, verbatim runs — measured +0.28% vs the 1% gate); local Verify: canonical g++ tree zero-warning, full `ctest` 89/89 (the new `profiler` entry + `api-real-tree` + `determinism-lint-real-tree` green after the manifest regeneration), `laige-run --prof-out` smoke (the summary line + the valid JSON report on disk); cross-tree builds + `tools/laige-include-lint` in the PR branch | | 2026-09-24 | M1-ALLOC-01 | `4564029` / PR #50 | Zero sim-loop allocation assertion (G-R1, PRD §9.3, §8.1 `sim_heap_allocs` target 0; PERF-003, FR-12.3; M1-ALLOC-01 scope, nothing else): the allocation watch (new public header `src/laige-core/include/laige/alloc_watch.h` + `src/laige-core/alloc_watch.cpp`) — a process-wide heap-allocation counter behind strong global `operator new`/`new[]` (+ nothrow, + sized deletes) in laige-core, compiled into every non-sanitizer tree (`LAIGE_ALLOC_WATCH=1` PUBLIC on laige-core; the sanitizer trees degrade to inline no-ops with `allocWatchLive()` false — the established fallback: the leak-free sanitizer run + the pool reservation delta, M0-CORE-02/05 precedent): the armed-window model (`allocWatchArm()` = three relaxed stores — first-site, count, armed flag; `allocWatchRead()` = two relaxed loads → `AllocWatchReading{allocs, firstSite}`; the single-owner window, the sim owner thread, CONC-001; first-site semantics: the allocating call's own return address — `__builtin_return_address(0)` on GCC/Clang, `_ReturnAddress()` on MSVC — evaluated in the operator-new frame, the first offender winning via a relaxed CAS that fails once recorded); the attribution contract: `laige::detail::LoggingAllocationGuard` — the logging facade's emit path (the `LAIGE_LOG` macro block + `Logger::record`) marks its own heap work (the field value strings, the rate-state, the sink's message formatting) so it is not attributed to the sim loop's G-R1 window — the engine's documented in-tick degradations (a G-R5 `budget_overrun`/`budget_critical`, a replay `record_failed`, a guardrail warn) still log (NFR-13.3 5-field grammar, rate-limited, actionable) and never trip G-R1, while any other in-tick allocation (a system's local `std::vector`, engine storage growth) still fails at its call site; the per-tick check (`GameLoop::runOneTick`, `#if !NDEBUG`, game_loop.cpp): arm BEFORE the tick body (the frame's `beginFrame` + one `runSystems` dispatch + the attached profiler + the `onTick` hook + the replay recorder), read AFTER a completed tick (`status.ok()` — a failed tick is not checked, the profiler's "a failed tick is not recorded" contract): a nonzero count logs one `alloc/sim_tick_allocation` Error event (fields `tick`/`allocs`/`site`, NFR-13.3) and then fails the debug assert (FR-12.3: actionable, never silent) — the standing hot-path guardrail for every later sim/render step (roadmap README §6, "Global invariants"); release builds compile the whole check out (CPP-012) — an allocating tick degrades through the already-logged pool accounting (pool overflow, pools.md) and the per-frame `simAllocs` delta (profiler.md) instead, never a crash; tests: the new `ZeroAlloc` suite + `zero_alloc` CTest entry (added to the TSan property list) over the shared `laige-sim_tests` executable — the 10k-entity M1-ECS-07 workload through the `GameLoop` (700 direct warm-up ticks bring every archetype to its high water BEFORE the window; then 10k ticks at 60 Hz on the synthetic clock — `kTickNs = 16666667` = ceil(10⁹/60), the game_loop_tests constant; the floor 16666666 drifts off the exact due count over 10k ticks) — per-tick window reads 0 allocs (the engine's own arm resets the watch each tick), the reservation delta 0, rows/entity invariants, the FNV-1a visit checksum, and the machine-greppable `zero-alloc window:` line; a scratch system with a deliberate `std::vector` fails the tick assert — proven in a forked SIGABRT child (POSIX; `GTEST_SKIP` on Windows), the release/sanitizer branch running 5 clean ticks (the Verify clause's deliberate-then-revert scratch kept as the standing negative test — the violation lives in the test TU, never in engine code); the watch's first-site capture checked directly; the test-side counter shim moved to laige-core (`tests/**/logging_alloc_counter.h` wraps the watch; the `LAIGE_ALLOC_COUNTER` test definition is gated on the same trees as `LAIGE_ALLOC_WATCH`, so the probes and the engine's assertion always agree); `HeadlessFramePathAllocatesNothing` (engine_tests) now reads per-tick window semantics (the engine's three one-shot setup allocations land before the first arm); docs in the same change: `docs/api/alloc_watch.md` (new — the window model, the per-tick assertion, the attribution contract, release builds, scope, cost, threading, misuse, example) + `game_loop.md` (the zero-allocation section + the Performance cost line) + `profiler.md` cross-ref + the docs/README index; `laige-api.json` regenerated (820 symbols from 25 headers, api-real-tree green); local Verify: `ctest -R zero_alloc` green, the full canonical ctest 92/92, and 92/92 on build-release/build-shared/build-asan/build-tsan/build-clang, zero warnings on every tree, the determinism + include lints OK; CI (observed via the GitHub API): the PR ci-pull.yml run 35988075605 on c9587fa green — all 8 jobs passed (linux-gcc/clang/asan+UBSan/tsan each ctest 92/92, the determinism check + source scan, the include-graph lint + dependency count, the public API manifest drift); macOS/Windows jobs skipped (label-gated, default-Linux P0 selection) | -| 2026-09-25 | M1-BENCH-01 | `—` | 10k-entity simulation-tick benchmark (PRD §8.1 ≤3.0 ms avg / ≤5 ms p99, ADR 0002, CORE-001; M1-BENCH-01 scope, nothing else): the `sim-tick` suite in `laige-bench` (`tools/bench/laige-bench.cpp` + CMake) — the PRD §8.1 "Simulation tick" workload: 10 000 entities at the 100% scene budget (capacity 10000, deterministic, seed `0x1F055EED` — the repo-wide test-seed convention), 2 000 dynamic bodies carrying the built-in `Position2D` plus the workload component `BenchVel` (a `{Vec2 v}`, `LAIGE_COMPONENT` + `LAIGE_DETERMINISM_SAFE` marks for both backends) and 8 000 bare entities; deterministic index-derived initial state (50×50 grid span, `vx ∈ -3..3`, `vy ∈ -5..5` units/tick — max drift 20 025 units over the 4 000-tick run, inside fpx16_16's ±32 768 Q16.16 range: no saturating-overflow edge in the measured window); two systems in registration order — `BenchMove` (1 ms declared budget; `pos += vel` through the active backend's `SimMath::add`, the ADR 0002 pinned op surface) and `BenchHash` (2 ms budget; the per-tick `World::stateHash(tick−1)` — the M1 determinism work made part of the measured tick) — driven at 60 Hz by the `GameLoop` on an EXACT synthetic clock (one due tick/frame: 16 666 667 ns — the M1-LOOP-01 exact integer due computation); one measured sample = one completed tick (beginFrame + the systems + the loop's bookkeeping; the presentation snapshot is excluded — ARCH-009); canonical shape `--runs=3000 --warmup=1000` (n=3000, histogram capacity 3000, no truncation); every measured tick in debug non-sanitizer builds passes the engine's OWN G-R1 per-tick zero-allocation assertion (M1-ALLOC-01 — an allocating tick aborts the run, so a PASS is a zero-alloc-asserted tick); tool changes (smallest complete change): `--math=` (the ADR 0002 backend selection; default fpx16; other suites ignore it), **repeatable** `--budget=` (one run gates EVERY named budget against the same histogram — each entry's `metric` picks its statistic; duplicate name = usage error; any FAIL exits 2), suite setup/teardown hooks (the stateless synthetic suite keeps nullptrs), and a compiler-id fix (Clang checked before GCC: Clang 22 defines `__GNUC__` and its `__VERSION__` is "Clang 22.1.8" — the old order mislabeled clang builds as "GCC Clang 22.1.8"; the id is now built from `__clang_major__/minor/patchlevel`, stable across Clang versions); CMake: `laige-bench` links `laige-sim` and gains `laige_apply_simmath_policy` (the TU carries deterministic sim math — ADR 0002 pinned flags), plus two CI gate entries `laige_bench_sim_tick_{fpx16,fp32}` (the canonical 3000/1000 shape, `--budget=sim_tick_avg --budget=sim_tick_p99`, `LAIGE_BUDGETS_PATH` = the repo root; EXCLUDED from the ASan/TSan trees — instrumentation inflates the tick's absolute cost: ASan+UBSan measures ≈1.6 ms/tick on this workload, 2.7× the canonical 0.59 ms, so a sanitizer measurement would measure the instrumentation, not the sim — the m1-profiler-cost precedent; those trees verify the workload's safety properties instead); the PRD §8.1 budget-regression gate is these entries inside every P0 job's full `ctest` (the CI perf lane, PRD §14 — no separate workflow step needed); **first recorded `measured` values in `budgets.json`** (the M0 convention's `0 = not yet measured` retired for the two entries — recorded as the WORSE of the two backends on the canonical Debug tree, since the budget covers both backends and the 10% regression band (methodology §5) stays conservative): `sim_tick_avg` (mean) **0.59283** ms, `sim_tick_p99` (p99) **0.618314** ms; **fourth baseline** `docs/benchmarks/baselines/m1-sim-tick.md` (full AGENTS §12 metadata, verbatim budget-checked runs + the ctest gate, the cross-tree/cross-compiler table, the interpretation — records that M1 measures the ECS-only slice of the §8.1 workload: kinematic movement + the state hash; the PRD's remaining tick content (physics, presentation) lands in M3/M4 and supersedes the baseline per methodology §4); docs updated in the same change (DOC-007): `docs/benchmarks/README.md` + `docs/benchmarks/baselines/README.md` (the stale "every entry still `measured: 0`" lines; the profiler-cost baseline entry that was never added), `docs/api/budget_harness.md` (the new CLI form, the sim-tick suite, the measured-field note, both suites' determinism); no new public API header — `laige-api.json` unchanged (api-real-tree green in the full-suite runs); local Verify (2026-09-25, AMD Ryzen 9 7950X3D / CachyOS): canonical Debug g++ 16.2.1 — fpx16_16 mean=0.59283 p99=0.618314, fp32_pinned mean=0.589047 p99=0.605951 — BOTH budgets PASS on BOTH backends (5.1× / 8.1× inside the 3.0/5.0 ms targets); Release g++ 0.0796/0.0839 (≈7.4× the Debug cost — the instrumented, unoptimized tree the gate measures); Clang 22.1.8 Debug 0.918488/0.939555 (1.5× GCC Debug — codegen difference, both far inside budget); shared-lib Debug 0.627088/0.651729; **the run's `final_hash` fingerprint is bit-identical across all four non-instrumented trees and both compilers, at both backends** (fpx16_16 `0x8aa6b855aa807b31`, fp32_pinned `0x34ff9f20dac09426` — the ADR 0002 cross-build bit-exactness demonstrated on the 4 000-tick workload); ASan (clang) run leak-free (exit 0, 600-tick safety shape); TSan `halt_on_error=1` race-free (exit 0) with one legitimate `system/budget_overrun` warn (BenchHash 2.04 ms vs its 2 ms budget under TSan's ≈4× slowdown — the engine's correct G-R5 observation, not a workload defect); full `ctest` green on all six local trees: 94/94 `build`/`build-clang`/`build-shared` (the 92 pre-existing + the two new gate entries), 93/93 `build-release` (one pre-existing Debug-only entry), 92/92 `build-asan`/`build-tsan` (the gate entries correctly absent there); zero new warnings under NFR-8.10; `tools/laige-include-lint` OK (47 source files, 1/10 deps), `tools/laige-determinism-lint` OK (28 sim files, 0 violations); CI: the PR's P0 jobs carry the gate entries in their full `ctest` (linux-gcc/clang default lane + the `ci:macos`/`ci:windows`-labeled windows-msvc/macos jobs + asan/tsan) — the M1-EXIT-01 gate records the final CI evidence (commit lands with the step's PR) | +| 2026-09-25 | M1-BENCH-01 | `—` | 10k-entity simulation-tick benchmark (PRD §8.1 ≤3.0 ms avg / ≤5 ms p99, ADR 0002, CORE-001; M1-BENCH-01 scope, nothing else): the `sim-tick` suite in `laige-bench` (`tools/bench/laige-bench.cpp` + CMake) — the PRD §8.1 "Simulation tick" workload: 10 000 entities at the 100% scene budget (capacity 10000, deterministic, seed `0x1F055EED` — the repo-wide test-seed convention), 2 000 dynamic bodies carrying the built-in `Position2D` plus the workload component `BenchVel` (a `{Vec2 v}`, `LAIGE_COMPONENT` + `LAIGE_DETERMINISM_SAFE` marks for both backends) and 8 000 bare entities; deterministic index-derived initial state (50×50 grid span — one body per cell of the first 2 000 cells, `x = (i/50)%50 − 25`, `y = i%50`; `vx ∈ -3..3`, `vy ∈ -5..5` units/tick — max coordinate magnitude 20 049 units over the 4 000-tick run, inside fpx16_16's ±32 768 Q16.16 range: no saturating-overflow edge in the measured window); two systems in registration order — `BenchMove` (1 ms declared budget; `pos += vel` through the active backend's `SimMath::add`, the ADR 0002 pinned op surface) and `BenchHash` (2 ms budget; the per-tick `World::stateHash(tick−1)` — the M1 determinism work made part of the measured tick) — driven at 60 Hz by the `GameLoop` on an EXACT synthetic clock (one due tick/frame: 16 666 667 ns — the M1-LOOP-01 exact integer due computation); one measured sample = one completed tick (beginFrame + the systems + the loop's bookkeeping; the presentation snapshot is excluded — ARCH-009); canonical shape `--runs=3000 --warmup=1000` (n=3000, histogram capacity 3000, no truncation); every measured tick in debug non-sanitizer builds passes the engine's OWN G-R1 per-tick zero-allocation assertion (M1-ALLOC-01 — an allocating tick aborts the run, so a PASS is a zero-alloc-asserted tick); tool changes (smallest complete change): `--math=` (the ADR 0002 backend selection; default fpx16; other suites ignore it), **repeatable** `--budget=` (one run gates EVERY named budget against the same histogram — each entry's `metric` picks its statistic; duplicate name = usage error; any FAIL exits 2), suite setup/teardown hooks (the stateless synthetic suite keeps nullptrs), and a compiler-id fix (Clang checked before GCC: Clang 22 defines `__GNUC__` and its `__VERSION__` is "Clang 22.1.8" — the old order mislabeled clang builds as "GCC Clang 22.1.8"; the id is now built from `__clang_major__/minor/patchlevel`, stable across Clang versions); CMake: `laige-bench` links `laige-sim` and gains `laige_apply_simmath_policy` (the TU carries deterministic sim math — ADR 0002 pinned flags), plus two CI gate entries `laige_bench_sim_tick_{fpx16,fp32}` (the canonical 3000/1000 shape, `--budget=sim_tick_avg --budget=sim_tick_p99`, `LAIGE_BUDGETS_PATH` = the repo root; EXCLUDED from the ASan/TSan trees — instrumentation inflates the tick's absolute cost: ASan+UBSan measures ≈1.6 ms/tick on this workload, 2.7× the canonical 0.59 ms, so a sanitizer measurement would measure the instrumentation, not the sim — the m1-profiler-cost precedent; those trees verify the workload's safety properties instead); the PRD §8.1 budget-regression gate is these entries inside every P0 job's full `ctest` (the CI perf lane, PRD §14 — no separate workflow step needed); **first recorded `measured` values in `budgets.json`** (the M0 convention's `0 = not yet measured` retired for the two entries — recorded as the WORSE of the two backends on the canonical Debug tree, since the budget covers both backends and the 10% regression band (methodology §5) stays conservative): `sim_tick_avg` (mean) **0.598405** ms, `sim_tick_p99` (p99) **0.615179** ms; **fourth baseline** `docs/benchmarks/baselines/m1-sim-tick.md` (full AGENTS §12 metadata, verbatim budget-checked runs + the ctest gate, the cross-tree/cross-compiler table, the interpretation — records that M1 measures the ECS-only slice of the §8.1 workload: kinematic movement + the state hash; the PRD's remaining tick content (physics, presentation) lands in M3/M4 and supersedes the baseline per methodology §4; the file also records the initial-state grid-shape fix folded into this PR — the first commit stacked the 2 000 bodies in one grid column instead of one-per-cell over the 50×50 span; re-measured on every tree (numbers moved <2%, both backends still PASS; the recorded numbers are the post-fix ones); docs updated in the same change (DOC-007): `docs/benchmarks/README.md` + `docs/benchmarks/baselines/README.md` (the stale "every entry still `measured: 0`" lines; the profiler-cost baseline entry that was never added), `docs/api/budget_harness.md` (the new CLI form, the sim-tick suite, the measured-field note, both suites' determinism); no new public API header — `laige-api.json` unchanged (api-real-tree green in the full-suite runs); local Verify (2026-09-25, AMD Ryzen 9 7950X3D / CachyOS): canonical Debug g++ 16.2.1 — fpx16_16 mean=0.598405 p99=0.606001, fp32_pinned mean=0.591438 p99=0.615179 — BOTH budgets PASS on BOTH backends (5.0× / 8.1× inside the 3.0/5.0 ms targets); Release g++ 0.0801446/0.08439 (≈7.5× the Debug cost — the instrumented, unoptimized tree the gate measures); Clang 22.1.8 Debug 0.913303/0.935979 (1.5× GCC Debug — codegen difference, both far inside budget); shared-lib Debug 0.623729/0.639315; **the run's `final_hash` fingerprint is bit-identical across all four non-instrumented trees and both compilers, at both backends** (fpx16_16 `0x7840644a24334cd0`, fp32_pinned `0x7a60d70232e448c7` — the ADR 0002 cross-build bit-exactness demonstrated on the 4 000-tick workload); ASan (clang) run leak-free (exit 0, 600-tick safety shape; ≈1.64 ms/tick — 2.7× the canonical, the instrumentation the gate excludes); TSan `halt_on_error=1` race-free (exit 0) with one legitimate `system/budget_overrun` warn (BenchHash 2.04 ms vs its 2 ms budget under TSan's ≈4× slowdown — the engine's correct G-R5 observation, not a workload defect); full `ctest` green on all six local trees: 94/94 `build`/`build-clang`/`build-shared` (the 92 pre-existing + the two new gate entries), 93/93 `build-release` (one pre-existing Debug-only entry), 92/92 `build-asan`/`build-tsan` (the gate entries correctly absent there); zero new warnings under NFR-8.10; `tools/laige-include-lint` OK (47 source files, 1/10 deps), `tools/laige-determinism-lint` OK (28 sim files, 0 violations); CI: the PR's P0 jobs carry the gate entries in their full `ctest` (linux-gcc/clang default lane + the `ci:macos`/`ci:windows`-labeled windows-msvc/macos jobs + asan/tsan) — the M1-EXIT-01 gate records the final CI evidence (commit lands with the step's PR) | --- diff --git a/tools/bench/laige-bench.cpp b/tools/bench/laige-bench.cpp index 460a9de..65fb3cf 100644 --- a/tools/bench/laige-bench.cpp +++ b/tools/bench/laige-bench.cpp @@ -164,12 +164,14 @@ constexpr std::uint64_t kSimTickSeed = 0x1F055EED; constexpr std::int64_t kSimTickNs = 16666667; // Initial-state shape (deterministic, index-derived — no RNG in the -// measured path): the dynamic bodies scatter over a 100-unit span -// (50×50 grid, one body per cell of the first 2000 cells) with a +// measured path): the dynamic bodies scatter over a 50×50 grid span +// (one body per cell of the first 2000 cells: x = col−25 ∈ −25..24, +// y = row ∈ 0..49) with a // small per-axis drift (vx ∈ -3..3, vy ∈ -5..5 world units per tick). -// Max drift over the full run (warmup + measured, 4 000 ticks) is -// 20 025 units — inside fpx16_16's ±32 768 Q16.16 range, so no -// saturating-overflow edge is ever touched in the measured window. +// Max coordinate magnitude over the full run (warmup + measured, +// 4 000 ticks) is 20 049 units (49 + 5·4000) — inside fpx16_16's +// ±32 768 Q16.16 range, so no saturating-overflow edge is ever +// touched in the measured window. constexpr std::int32_t kSimTickGridSpan = 50; constexpr std::int32_t kSimTickVelXMax = 3; constexpr std::int32_t kSimTickVelYMax = 5; @@ -323,8 +325,10 @@ bool simTickSetup(const Config& cfg) { } const laige::Entity entity = std::move(e).takeValue(); if (i < kSimTickDynamic) { + // One body per grid cell, column-major over the 50×50 span: + // i = 50·col + row — the first 2 000 cells (40 columns of 50). const std::int32_t x = - static_cast((i / kSimTickDynamic) % + static_cast((i / kSimTickGridSpan) % kSimTickGridSpan) - kSimTickGridSpan / 2; const std::int32_t y = static_cast(i % kSimTickGridSpan); From f77300c22d1b0ad70e0f789d1316071b5374b0ba Mon Sep 17 00:00:00 2001 From: Pascal Severin Date: Fri, 25 Sep 2026 17:45:22 +0200 Subject: [PATCH 4/5] CI trigger: windows-msvc P0 job (label-gated matrix, M1-EXIT-01 all-OS evidence) The PR-lane label protocol runs one P0 OS family per run: no labels = linux-gcc/clang/asan/tsan, ci:windows (no ci:macos) = windows-msvc, ci:macos = macos-arm64/intel. The M1-EXIT-01 gate requires replay bit-exactness evidence on every P0 OS job for this PR's code; the Linux lane ran on 36152843429 (c2287c1), the macOS pair on 36155154836 (this commit's parent), this empty commit selects the Windows lane. From 2dbe637d78f0d355d9677708f7b8cb91679a5625 Mon Sep 17 00:00:00 2001 From: Pascal Severin Date: Fri, 25 Sep 2026 17:50:07 +0200 Subject: [PATCH 5/5] [M1-EXIT-01] Record the M1 exit gate evidence MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The three gate items, with their CI evidence links: (1) sim-tick budget green on both SimMath backends — the baseline report + PR-lane runs 36152843429 (linux-gcc/clang full ctest 94/94 with the two budget-gate entries) and 36155154836 (macOS arm64/Intel full ctest, same entries green on AppleClang); the CI reference machine (ubuntu-24.04) is the gate. (2) replay bit-exact on CI across all P0 OS jobs and both backends — the determinism matrix + merge-lane run 36037080146 (all 7 P0 jobs + detcheck on master) and this PR's label-gated P0 matrix: Linux run 36152843429, macOS arm64/Intel run 36155154836, windows-msvc run 36156322585 — every job green, including hello_baseline_fpx/fp32 in each OS job's ctest and the detcheck job (both backends). (3) zero-alloc assertion green on the 10k workload — the zero_alloc ctest entry green in every P0 job's full ctest in the runs above, and the M1-BENCH-01 measurement itself zero-alloc-asserted per tick (G-R1). Deferred P1 items: none — every M1 step is P0. Progress Board: M1 25/25 complete (total 47/193), set in the prior commit; the M1-BENCH-01 changelog line's stale 1.6/0.59 ms ASan mention corrected to the post-fix 1.64/0.60 ms in the same change. --- roadmap/M1-heartbeat.md | 40 ++++++++++++++++++++++++++++++++++++++++ roadmap/README.md | 3 ++- 2 files changed, 42 insertions(+), 1 deletion(-) diff --git a/roadmap/M1-heartbeat.md b/roadmap/M1-heartbeat.md index 85c3b46..4a4df0f 100644 --- a/roadmap/M1-heartbeat.md +++ b/roadmap/M1-heartbeat.md @@ -290,4 +290,44 @@ zero-allocation property (M1-ALLOC-01 enforces it once it exists; before that, A - Confirm and record: (1) `sim-tick` budget green on both SimMath backends (link report), (2) replay bit-exact on CI across all P0 OS jobs and both backends (link run), (3) zero-alloc assertion green on the 10k workload (link test log). - Update Progress Board; note any deferred P1 items (there should be none in M1 — everything here is P0). - **Verify:** all three evidence links present; no open M1 step. + - **Recorded (2026-09-25):** gate items confirmed and recorded: + (1) **`sim-tick` budget green on both SimMath backends** — the + report [docs/benchmarks/baselines/m1-sim-tick.md](../docs/benchmarks/baselines/m1-sim-tick.md) + (canonical Debug, both backends PASS: `fixed_point_16_16` mean + 0.598405 / p99 0.606001 ms, `float_pinned_32` mean 0.591438 / + p99 0.615179 ms, against the 3.0 / 5.0 ms targets) — and on CI + (the CI reference machine, ubuntu-24.04, is the gate): PR-lane run + [36152843429](https://github.com/offdev/laige-cpp/actions/runs/36152843429) + (linux-gcc + linux-clang full `ctest` 94/94, including the two + budget-gate entries `laige_bench_sim_tick_fpx16` / + `laige_bench_sim_tick_fp32` — both budgets, both backends, exit 0) + and run [36155154836](https://github.com/offdev/laige-cpp/actions/runs/36155154836) + (macOS arm64 + Intel full `ctest`, the same two gate entries green + on AppleClang). + (2) **Replay bit-exact on CI across all P0 OS jobs and both + backends** — the determinism report + [docs/benchmarks/determinism-matrix.md](../docs/benchmarks/determinism-matrix.md) + (every P0 OS job asserts per-tick identity of the hello baseline + for both backends — `hello_baseline_fpx` / `hello_baseline_fp32` — + and the always-on detcheck job compares the cross-build pairs on + both backends) — CI evidence: merge-lane run + [36037080146](https://github.com/offdev/laige-cpp/actions/runs/36037080146) + (master, 2026-09-24: all 7 P0 jobs + detcheck green) and this PR's + label-gated P0 matrix — Linux: run 36152843429, macOS arm64 / + Intel: run 36155154836, Windows x64: run + [36156322585](https://github.com/offdev/laige-cpp/actions/runs/36156322585) + — every job green, including the two hello-baseline tests in each + OS job's full `ctest` and the detcheck job (both backends). + (3) **Zero-alloc assertion green on the 10k workload** — the + `zero_alloc` ctest entry (M1-ALLOC-01: the M1-ECS-07 10 000-entity + workload driven through the game loop, asserting 0 allocations per + tick in the debug trees) green in every P0 job's full `ctest` in + the runs above (test logs in each run's job logs; the ASan/TSan + lanes additionally archive the sanitizer reports) — and the + M1-BENCH-01 measurement itself is zero-alloc-asserted: every tick + of both backends (4 000 each — warm-up and measured) passed the + engine's G-R1 per-tick allocation assertion (an allocating tick + aborts the run). + Deferred P1 items: none — every M1 step is P0. + Progress Board: M1 25/25 complete (total 47/193). - **Size:** docs only diff --git a/roadmap/README.md b/roadmap/README.md index efe3be9..da28c36 100644 --- a/roadmap/README.md +++ b/roadmap/README.md @@ -217,7 +217,8 @@ One line per completed (or split/renumbered) step. | 2026-09-23 | M1-PROF-02 | `0c7cf5c` / PR #49 | Frame graph / budget report (FR-11.2, PRD §9.1 S-6, §9.3 G-R5; M1-PROF-02 scope, nothing else): the `FrameBudgetRecorder` fixed 32-frame ring (`src/laige-sim/include/laige/sim/frame_budget.h` + `frame_budget.cpp` — `FrameBudgetRecord` per completed frame: the 0-based frame index, the tick delta, the frame's sim work ms, the pool-reservations sim-alloc delta, the G-R5 `budget_overrun`/`budget_critical` event deltas; `recordFrame` O(1) allocation-free hot path, `at(i)` oldest-first over the wrapping ring, `totalFrames()` keeps counting past the window — no silent truncation, CORE-008) and `buildFrameBudgetReport` (cold format pass, the AGENTS §12 field format): every DECLARED budget measured vs declared with a pass/flag — each system's declared `SystemDef::budgetMs` (fpx16_16, exact — ADR 0002) vs its M1-SYS-03 rolling window's **p99** (the G-R5 sustained-overrun signal; single recovered overruns stay visible in the per-frame records + the `warns`/`errors` counters), `sim_tick_avg`/`sim_tick_p99` (budgets.json, M0-CORE-08) vs the `Profiler`'s tick window (mean/p99), `sim_heap_allocs` (the PRD §8.1 hard-zero budget) vs the per-frame sim-alloc deltas (max); the M0-CORE-08 `budgetCheck` blocks embedded verbatim; a missing entry is a loud `NO_ENTRY`, an empty window a loud `NO_SAMPLES` (a zero-tick run is never silent), the over-budget systems list ascending id, and `overall=PASS|FAIL` (a broken harness folds to FAIL — CORE-008); the ENGINE wiring (always-on per-frame accumulation — two O(1) reads + two O(systemCount) G-R5 counter passes + one O(1) ring write per frame, no allocation, no logging, PERF-003/LOG-003; the run-setup allocation count stays exactly three — `HeadlessFramePathAllocatesNothing` still pins it — the ring is created in `Engine::create`): `startBudgetReport(budgetsPath, lastNFrames)` (EVERY build — diagnostics, not replay state; loads the table now, builds + CACHES the report at the run's end on EVERY path — readable after the shutdown, the `profileStats()` precedent; a budget FAIL never fails the run — CORE-002; errors: stopped/empty-path/double-start `InvalidArgument`, load `IoError`/`MalformedInput` + `budget/report_*` structured events), `budgetReportRequested()`, `lastBudgetReport()`; the CLI surface (`tools/run/laige-run.cpp`): `--budget-report [N]` (1..32; omitted = all retained; printed to stdout after the profile summary line, even on a failed run), `--budgets ` (flag → `LAIGE_BUDGETS_PATH` env → `budgets.json` CWD — the laige-bench resolution order), `--fail-on-budget` (**exit 3** when the run COMPLETED but `overall=FAIL` — the PRD §8.1 CI gate; run failure stays exit 1, start/load failure exit 2); 14-test `budget_report` CTest entry (`tests/laige-sim/budget_report_tests.cpp`: the ring's newest-frames-oldest-first semantics, the SYNTHETIC OVER-BUDGET SYSTEM with correct numbers — the 2 ms burn vs a 0.1 ms fpx16_16 budget (20× — both G-R5 multipliers exceeded on the nominal floor, so the exact `runs=10 warns=10 errors=10` and p99 ≥ 2 assertions are preemption-tolerant): the declared budget echoed, the `over_budget:` line, `overall=FAIL`, the per-frame FAIL invariants — the engine's cached-after-shutdown report, the G-R5 events folded into the per-frame records (rate-limiting-off capture sink: 10 warns + 10 criticals exact), the healthy-world PASS, the loud NO_ENTRY/NO_SAMPLES, the start validation (double/empty/stopped/unreadable/malformed), and the record path's zero-allocation (`budget-report-recorder-zeroalloc frames=1000 allocs=0`, non-sanitizer trees)); `laige_run_budget` CLI smoke (every P0 OS job: 30-tick run, `--budget-report 4 --budgets /budgets.json --fail-on-budget`, passes iff `overall=PASS` + exit 0); sample report committed as `tests/laige-sim/fixtures/budget_report_sample.txt` (a sample, NOT a golden — the report carries wall-clock values); G-R8 exception markers on the raw `double` tokens (wall-clock diagnostics — never enter sim state, hashes, or replays, ARCH-009; `tools/laige-determinism-lint` green); `laige-api.json` regenerated (24 headers; +`FrameBudgetRecord`/`kFrameBudgetWindow`/`FrameBudgetRecorder`/`FrameBudgetReportOptions`/`FrameBudgetReport`/`buildFrameBudgetReport` + the three `Engine` methods + `Profiler::tickWindow`); docs in the same change: `docs/api/frame_budget.md` (new — the declared budgets, the report format + pass/flag semantics, the Engine surface, the CLI + exit 3, Performance, determinism, Testing/CI) + engine.md (CLI flags, exit 3, per-frame cost, misuse, Testing) + profiler.md cross-ref + the docs/README + sim README indexes; local Verify: canonical g++ Debug tree zero-warning, `ctest -R budget_report` green (14/14), `ctest -R laige_run_budget` green, `ctest -R profiler`/`engine`/`system_timing` green (the zero-allocation pin unchanged), both lints OK | | 2026-09-21 | M1-PROF-01 | `a811297` / PR #47 | Always-on profiler counters (FR-11.1, DBG-008; M1-PROF-01 scope, nothing else): the `Profiler` core (`src/laige-sim/include/laige/sim/profiler.h` + `profiler.cpp`) — always-on, fixed-storage, no allocation after init: tick/frame time rolling windows (512/256 samples, M0-CORE-08 `Histogram`; record is O(1) allocation-free, drops the OLDEST, the since-construction counters keep counting), draw calls / texture binds / net bytes counters (0 in headless M1 — the fields exist per FR-11.1, the render/network subsystems feed them in M2/M3), and the cold `snapshot()` (own counters; `snapshot(world)` adds the world-pulled fields — entity total/alive/capacity, sim alloc count = `World::archetypeStats().totalReservations` (the M1-ECS-03 pool accounting, target 0), system count — read COLD, never copied; per-system windows stay in `World`, pulled only in the report); move-only (moved-from = stopped: records no-op, empty snapshot); `setEnabled`/`enabled` (disabled = one branch, no recording); report surface: `formatProfileSummaryLine` (the CLI one-liner), `formatProfileText` / `formatProfileJson` (version-1 schema: counters, tick/frame windows (n==0 → `n=0` text / JSON `null`, never NaN), world fields, per-system entries with the M1-SYS-03 window stats), `writeProfile` (truncating write, no partial file on failure, `Result` — `IoError`); wiring: `GameLoop::Options::profiler` (non-owning) — `runOneTick` times the tick body (the frame's `beginFrame` + one `runSystems` dispatch) with the M0-CORE-08 `TimeIt` and records on SUCCESS only (a failed tick is neither counted nor recorded), null/disabled = one branch; the engine owns the profiler (`Engine::create` constructs it — engine setup, not run setup, so the "exactly three one-shot allocations per run" claim stays true; the run loop adds the frame feed — two clock reads + one ring write per frame, excluding the pacing sleep, first frame not recorded, failed frames not recorded; `shutdown()` releases it in the pools step, abandoning a started-but-unfinalized report with `profiler/report_aborted`); the per-run report: `startProfileReport(path)` (EVERY build — diagnostics, not replay state, no `NDEBUG` gate), written at the END of the run on EVERY path (a zero-tick run writes a zero-tick report, CORE-008), a write failure does NOT fail the run (sticky `profileReportStatus()`, `profiler/report_write_failed` Error), `profileStats()` = the last run's cached snapshot (the world is released in shutdown); `laige-run --prof-out ` (start failure exits 2; a write failure leaves the run `status=ok` and exits 2) + the always-printed `laige-run profile: …` one-line summary (the byte-stable `status=ok` line untouched — the summary is a separate stdout line); structured events (subsystem `profiler`, NFR-13.3 5-field grammar): `report_started` (Info), `report_written` (Info), `report_write_failed` (Error), `report_aborted` / `report_already_started` / `report_path_invalid` (Warn); G-R8 exception markers on the raw `double` tokens (wall-clock diagnostic — never enters sim state, hashes, or replays, ARCH-009); 26-test `profiler` CTest entry (`tests/laige-sim/profiler_tests.cpp`: exact percentiles 1..100 → p50=50/p95=95/p99=99/mean=50.5, rollover cap, independent windows, zero-capacity drop, adders, disabled no-op + preserved state, moved-from stop, cold snapshot, per-completed-tick timing (failed tick unrecorded), the engine's per-run cache + report (written on every run path, version-1 JSON parseable, double-start/empty-path/stopped-engine rejections, write failure sticky without failing the run, pre-run shutdown abandonment with no file on disk), the greppable text form, the record path's zero-allocation (`profiler-zeroalloc ticks=1000 allocs=0`, non-sanitizer trees), and the enabled-cost gate ON vs OFF over 10k-entity ticks ≤ 1% (`profiler-cost on_p50=0.557288 off_p50=0.555746 overhead_pct=0.277465`, best-of-2 per arm, non-sanitizer trees — the LAIGE_ALLOC_COUNTER gate: sanitizer instrumentation inflates the fixed per-tick cost, 1.46% on the ASan tree)); docs in the same change: `docs/api/profiler.md` (new) + engine.md / game_loop.md updates + the docs/README, sim README, debugging README, tools README indexes; baseline `docs/benchmarks/baselines/m1-profiler-cost.md` (full AGENTS §12 metadata, verbatim runs — measured +0.28% vs the 1% gate); local Verify: canonical g++ tree zero-warning, full `ctest` 89/89 (the new `profiler` entry + `api-real-tree` + `determinism-lint-real-tree` green after the manifest regeneration), `laige-run --prof-out` smoke (the summary line + the valid JSON report on disk); cross-tree builds + `tools/laige-include-lint` in the PR branch | | 2026-09-24 | M1-ALLOC-01 | `4564029` / PR #50 | Zero sim-loop allocation assertion (G-R1, PRD §9.3, §8.1 `sim_heap_allocs` target 0; PERF-003, FR-12.3; M1-ALLOC-01 scope, nothing else): the allocation watch (new public header `src/laige-core/include/laige/alloc_watch.h` + `src/laige-core/alloc_watch.cpp`) — a process-wide heap-allocation counter behind strong global `operator new`/`new[]` (+ nothrow, + sized deletes) in laige-core, compiled into every non-sanitizer tree (`LAIGE_ALLOC_WATCH=1` PUBLIC on laige-core; the sanitizer trees degrade to inline no-ops with `allocWatchLive()` false — the established fallback: the leak-free sanitizer run + the pool reservation delta, M0-CORE-02/05 precedent): the armed-window model (`allocWatchArm()` = three relaxed stores — first-site, count, armed flag; `allocWatchRead()` = two relaxed loads → `AllocWatchReading{allocs, firstSite}`; the single-owner window, the sim owner thread, CONC-001; first-site semantics: the allocating call's own return address — `__builtin_return_address(0)` on GCC/Clang, `_ReturnAddress()` on MSVC — evaluated in the operator-new frame, the first offender winning via a relaxed CAS that fails once recorded); the attribution contract: `laige::detail::LoggingAllocationGuard` — the logging facade's emit path (the `LAIGE_LOG` macro block + `Logger::record`) marks its own heap work (the field value strings, the rate-state, the sink's message formatting) so it is not attributed to the sim loop's G-R1 window — the engine's documented in-tick degradations (a G-R5 `budget_overrun`/`budget_critical`, a replay `record_failed`, a guardrail warn) still log (NFR-13.3 5-field grammar, rate-limited, actionable) and never trip G-R1, while any other in-tick allocation (a system's local `std::vector`, engine storage growth) still fails at its call site; the per-tick check (`GameLoop::runOneTick`, `#if !NDEBUG`, game_loop.cpp): arm BEFORE the tick body (the frame's `beginFrame` + one `runSystems` dispatch + the attached profiler + the `onTick` hook + the replay recorder), read AFTER a completed tick (`status.ok()` — a failed tick is not checked, the profiler's "a failed tick is not recorded" contract): a nonzero count logs one `alloc/sim_tick_allocation` Error event (fields `tick`/`allocs`/`site`, NFR-13.3) and then fails the debug assert (FR-12.3: actionable, never silent) — the standing hot-path guardrail for every later sim/render step (roadmap README §6, "Global invariants"); release builds compile the whole check out (CPP-012) — an allocating tick degrades through the already-logged pool accounting (pool overflow, pools.md) and the per-frame `simAllocs` delta (profiler.md) instead, never a crash; tests: the new `ZeroAlloc` suite + `zero_alloc` CTest entry (added to the TSan property list) over the shared `laige-sim_tests` executable — the 10k-entity M1-ECS-07 workload through the `GameLoop` (700 direct warm-up ticks bring every archetype to its high water BEFORE the window; then 10k ticks at 60 Hz on the synthetic clock — `kTickNs = 16666667` = ceil(10⁹/60), the game_loop_tests constant; the floor 16666666 drifts off the exact due count over 10k ticks) — per-tick window reads 0 allocs (the engine's own arm resets the watch each tick), the reservation delta 0, rows/entity invariants, the FNV-1a visit checksum, and the machine-greppable `zero-alloc window:` line; a scratch system with a deliberate `std::vector` fails the tick assert — proven in a forked SIGABRT child (POSIX; `GTEST_SKIP` on Windows), the release/sanitizer branch running 5 clean ticks (the Verify clause's deliberate-then-revert scratch kept as the standing negative test — the violation lives in the test TU, never in engine code); the watch's first-site capture checked directly; the test-side counter shim moved to laige-core (`tests/**/logging_alloc_counter.h` wraps the watch; the `LAIGE_ALLOC_COUNTER` test definition is gated on the same trees as `LAIGE_ALLOC_WATCH`, so the probes and the engine's assertion always agree); `HeadlessFramePathAllocatesNothing` (engine_tests) now reads per-tick window semantics (the engine's three one-shot setup allocations land before the first arm); docs in the same change: `docs/api/alloc_watch.md` (new — the window model, the per-tick assertion, the attribution contract, release builds, scope, cost, threading, misuse, example) + `game_loop.md` (the zero-allocation section + the Performance cost line) + `profiler.md` cross-ref + the docs/README index; `laige-api.json` regenerated (820 symbols from 25 headers, api-real-tree green); local Verify: `ctest -R zero_alloc` green, the full canonical ctest 92/92, and 92/92 on build-release/build-shared/build-asan/build-tsan/build-clang, zero warnings on every tree, the determinism + include lints OK; CI (observed via the GitHub API): the PR ci-pull.yml run 35988075605 on c9587fa green — all 8 jobs passed (linux-gcc/clang/asan+UBSan/tsan each ctest 92/92, the determinism check + source scan, the include-graph lint + dependency count, the public API manifest drift); macOS/Windows jobs skipped (label-gated, default-Linux P0 selection) | -| 2026-09-25 | M1-BENCH-01 | `—` | 10k-entity simulation-tick benchmark (PRD §8.1 ≤3.0 ms avg / ≤5 ms p99, ADR 0002, CORE-001; M1-BENCH-01 scope, nothing else): the `sim-tick` suite in `laige-bench` (`tools/bench/laige-bench.cpp` + CMake) — the PRD §8.1 "Simulation tick" workload: 10 000 entities at the 100% scene budget (capacity 10000, deterministic, seed `0x1F055EED` — the repo-wide test-seed convention), 2 000 dynamic bodies carrying the built-in `Position2D` plus the workload component `BenchVel` (a `{Vec2 v}`, `LAIGE_COMPONENT` + `LAIGE_DETERMINISM_SAFE` marks for both backends) and 8 000 bare entities; deterministic index-derived initial state (50×50 grid span — one body per cell of the first 2 000 cells, `x = (i/50)%50 − 25`, `y = i%50`; `vx ∈ -3..3`, `vy ∈ -5..5` units/tick — max coordinate magnitude 20 049 units over the 4 000-tick run, inside fpx16_16's ±32 768 Q16.16 range: no saturating-overflow edge in the measured window); two systems in registration order — `BenchMove` (1 ms declared budget; `pos += vel` through the active backend's `SimMath::add`, the ADR 0002 pinned op surface) and `BenchHash` (2 ms budget; the per-tick `World::stateHash(tick−1)` — the M1 determinism work made part of the measured tick) — driven at 60 Hz by the `GameLoop` on an EXACT synthetic clock (one due tick/frame: 16 666 667 ns — the M1-LOOP-01 exact integer due computation); one measured sample = one completed tick (beginFrame + the systems + the loop's bookkeeping; the presentation snapshot is excluded — ARCH-009); canonical shape `--runs=3000 --warmup=1000` (n=3000, histogram capacity 3000, no truncation); every measured tick in debug non-sanitizer builds passes the engine's OWN G-R1 per-tick zero-allocation assertion (M1-ALLOC-01 — an allocating tick aborts the run, so a PASS is a zero-alloc-asserted tick); tool changes (smallest complete change): `--math=` (the ADR 0002 backend selection; default fpx16; other suites ignore it), **repeatable** `--budget=` (one run gates EVERY named budget against the same histogram — each entry's `metric` picks its statistic; duplicate name = usage error; any FAIL exits 2), suite setup/teardown hooks (the stateless synthetic suite keeps nullptrs), and a compiler-id fix (Clang checked before GCC: Clang 22 defines `__GNUC__` and its `__VERSION__` is "Clang 22.1.8" — the old order mislabeled clang builds as "GCC Clang 22.1.8"; the id is now built from `__clang_major__/minor/patchlevel`, stable across Clang versions); CMake: `laige-bench` links `laige-sim` and gains `laige_apply_simmath_policy` (the TU carries deterministic sim math — ADR 0002 pinned flags), plus two CI gate entries `laige_bench_sim_tick_{fpx16,fp32}` (the canonical 3000/1000 shape, `--budget=sim_tick_avg --budget=sim_tick_p99`, `LAIGE_BUDGETS_PATH` = the repo root; EXCLUDED from the ASan/TSan trees — instrumentation inflates the tick's absolute cost: ASan+UBSan measures ≈1.6 ms/tick on this workload, 2.7× the canonical 0.59 ms, so a sanitizer measurement would measure the instrumentation, not the sim — the m1-profiler-cost precedent; those trees verify the workload's safety properties instead); the PRD §8.1 budget-regression gate is these entries inside every P0 job's full `ctest` (the CI perf lane, PRD §14 — no separate workflow step needed); **first recorded `measured` values in `budgets.json`** (the M0 convention's `0 = not yet measured` retired for the two entries — recorded as the WORSE of the two backends on the canonical Debug tree, since the budget covers both backends and the 10% regression band (methodology §5) stays conservative): `sim_tick_avg` (mean) **0.598405** ms, `sim_tick_p99` (p99) **0.615179** ms; **fourth baseline** `docs/benchmarks/baselines/m1-sim-tick.md` (full AGENTS §12 metadata, verbatim budget-checked runs + the ctest gate, the cross-tree/cross-compiler table, the interpretation — records that M1 measures the ECS-only slice of the §8.1 workload: kinematic movement + the state hash; the PRD's remaining tick content (physics, presentation) lands in M3/M4 and supersedes the baseline per methodology §4; the file also records the initial-state grid-shape fix folded into this PR — the first commit stacked the 2 000 bodies in one grid column instead of one-per-cell over the 50×50 span; re-measured on every tree (numbers moved <2%, both backends still PASS; the recorded numbers are the post-fix ones); docs updated in the same change (DOC-007): `docs/benchmarks/README.md` + `docs/benchmarks/baselines/README.md` (the stale "every entry still `measured: 0`" lines; the profiler-cost baseline entry that was never added), `docs/api/budget_harness.md` (the new CLI form, the sim-tick suite, the measured-field note, both suites' determinism); no new public API header — `laige-api.json` unchanged (api-real-tree green in the full-suite runs); local Verify (2026-09-25, AMD Ryzen 9 7950X3D / CachyOS): canonical Debug g++ 16.2.1 — fpx16_16 mean=0.598405 p99=0.606001, fp32_pinned mean=0.591438 p99=0.615179 — BOTH budgets PASS on BOTH backends (5.0× / 8.1× inside the 3.0/5.0 ms targets); Release g++ 0.0801446/0.08439 (≈7.5× the Debug cost — the instrumented, unoptimized tree the gate measures); Clang 22.1.8 Debug 0.913303/0.935979 (1.5× GCC Debug — codegen difference, both far inside budget); shared-lib Debug 0.623729/0.639315; **the run's `final_hash` fingerprint is bit-identical across all four non-instrumented trees and both compilers, at both backends** (fpx16_16 `0x7840644a24334cd0`, fp32_pinned `0x7a60d70232e448c7` — the ADR 0002 cross-build bit-exactness demonstrated on the 4 000-tick workload); ASan (clang) run leak-free (exit 0, 600-tick safety shape; ≈1.64 ms/tick — 2.7× the canonical, the instrumentation the gate excludes); TSan `halt_on_error=1` race-free (exit 0) with one legitimate `system/budget_overrun` warn (BenchHash 2.04 ms vs its 2 ms budget under TSan's ≈4× slowdown — the engine's correct G-R5 observation, not a workload defect); full `ctest` green on all six local trees: 94/94 `build`/`build-clang`/`build-shared` (the 92 pre-existing + the two new gate entries), 93/93 `build-release` (one pre-existing Debug-only entry), 92/92 `build-asan`/`build-tsan` (the gate entries correctly absent there); zero new warnings under NFR-8.10; `tools/laige-include-lint` OK (47 source files, 1/10 deps), `tools/laige-determinism-lint` OK (28 sim files, 0 violations); CI: the PR's P0 jobs carry the gate entries in their full `ctest` (linux-gcc/clang default lane + the `ci:macos`/`ci:windows`-labeled windows-msvc/macos jobs + asan/tsan) — the M1-EXIT-01 gate records the final CI evidence (commit lands with the step's PR) | +| 2026-09-25 | M1-BENCH-01 | `—` | 10k-entity simulation-tick benchmark (PRD §8.1 ≤3.0 ms avg / ≤5 ms p99, ADR 0002, CORE-001; M1-BENCH-01 scope, nothing else): the `sim-tick` suite in `laige-bench` (`tools/bench/laige-bench.cpp` + CMake) — the PRD §8.1 "Simulation tick" workload: 10 000 entities at the 100% scene budget (capacity 10000, deterministic, seed `0x1F055EED` — the repo-wide test-seed convention), 2 000 dynamic bodies carrying the built-in `Position2D` plus the workload component `BenchVel` (a `{Vec2 v}`, `LAIGE_COMPONENT` + `LAIGE_DETERMINISM_SAFE` marks for both backends) and 8 000 bare entities; deterministic index-derived initial state (50×50 grid span — one body per cell of the first 2 000 cells, `x = (i/50)%50 − 25`, `y = i%50`; `vx ∈ -3..3`, `vy ∈ -5..5` units/tick — max coordinate magnitude 20 049 units over the 4 000-tick run, inside fpx16_16's ±32 768 Q16.16 range: no saturating-overflow edge in the measured window); two systems in registration order — `BenchMove` (1 ms declared budget; `pos += vel` through the active backend's `SimMath::add`, the ADR 0002 pinned op surface) and `BenchHash` (2 ms budget; the per-tick `World::stateHash(tick−1)` — the M1 determinism work made part of the measured tick) — driven at 60 Hz by the `GameLoop` on an EXACT synthetic clock (one due tick/frame: 16 666 667 ns — the M1-LOOP-01 exact integer due computation); one measured sample = one completed tick (beginFrame + the systems + the loop's bookkeeping; the presentation snapshot is excluded — ARCH-009); canonical shape `--runs=3000 --warmup=1000` (n=3000, histogram capacity 3000, no truncation); every measured tick in debug non-sanitizer builds passes the engine's OWN G-R1 per-tick zero-allocation assertion (M1-ALLOC-01 — an allocating tick aborts the run, so a PASS is a zero-alloc-asserted tick); tool changes (smallest complete change): `--math=` (the ADR 0002 backend selection; default fpx16; other suites ignore it), **repeatable** `--budget=` (one run gates EVERY named budget against the same histogram — each entry's `metric` picks its statistic; duplicate name = usage error; any FAIL exits 2), suite setup/teardown hooks (the stateless synthetic suite keeps nullptrs), and a compiler-id fix (Clang checked before GCC: Clang 22 defines `__GNUC__` and its `__VERSION__` is "Clang 22.1.8" — the old order mislabeled clang builds as "GCC Clang 22.1.8"; the id is now built from `__clang_major__/minor/patchlevel`, stable across Clang versions); CMake: `laige-bench` links `laige-sim` and gains `laige_apply_simmath_policy` (the TU carries deterministic sim math — ADR 0002 pinned flags), plus two CI gate entries `laige_bench_sim_tick_{fpx16,fp32}` (the canonical 3000/1000 shape, `--budget=sim_tick_avg --budget=sim_tick_p99`, `LAIGE_BUDGETS_PATH` = the repo root; EXCLUDED from the ASan/TSan trees — instrumentation inflates the tick's absolute cost: ASan+UBSan measures ≈1.64 ms/tick on this workload, 2.7× the canonical 0.60 ms, so a sanitizer measurement would measure the instrumentation, not the sim — the m1-profiler-cost precedent; those trees verify the workload's safety properties instead); the PRD §8.1 budget-regression gate is these entries inside every P0 job's full `ctest` (the CI perf lane, PRD §14 — no separate workflow step needed); **first recorded `measured` values in `budgets.json`** (the M0 convention's `0 = not yet measured` retired for the two entries — recorded as the WORSE of the two backends on the canonical Debug tree, since the budget covers both backends and the 10% regression band (methodology §5) stays conservative): `sim_tick_avg` (mean) **0.598405** ms, `sim_tick_p99` (p99) **0.615179** ms; **fourth baseline** `docs/benchmarks/baselines/m1-sim-tick.md` (full AGENTS §12 metadata, verbatim budget-checked runs + the ctest gate, the cross-tree/cross-compiler table, the interpretation — records that M1 measures the ECS-only slice of the §8.1 workload: kinematic movement + the state hash; the PRD's remaining tick content (physics, presentation) lands in M3/M4 and supersedes the baseline per methodology §4; the file also records the initial-state grid-shape fix folded into this PR — the first commit stacked the 2 000 bodies in one grid column instead of one-per-cell over the 50×50 span; re-measured on every tree (numbers moved <2%, both backends still PASS; the recorded numbers are the post-fix ones); docs updated in the same change (DOC-007): `docs/benchmarks/README.md` + `docs/benchmarks/baselines/README.md` (the stale "every entry still `measured: 0`" lines; the profiler-cost baseline entry that was never added), `docs/api/budget_harness.md` (the new CLI form, the sim-tick suite, the measured-field note, both suites' determinism); no new public API header — `laige-api.json` unchanged (api-real-tree green in the full-suite runs); local Verify (2026-09-25, AMD Ryzen 9 7950X3D / CachyOS): canonical Debug g++ 16.2.1 — fpx16_16 mean=0.598405 p99=0.606001, fp32_pinned mean=0.591438 p99=0.615179 — BOTH budgets PASS on BOTH backends (5.0× / 8.1× inside the 3.0/5.0 ms targets); Release g++ 0.0801446/0.08439 (≈7.5× the Debug cost — the instrumented, unoptimized tree the gate measures); Clang 22.1.8 Debug 0.913303/0.935979 (1.5× GCC Debug — codegen difference, both far inside budget); shared-lib Debug 0.623729/0.639315; **the run's `final_hash` fingerprint is bit-identical across all four non-instrumented trees and both compilers, at both backends** (fpx16_16 `0x7840644a24334cd0`, fp32_pinned `0x7a60d70232e448c7` — the ADR 0002 cross-build bit-exactness demonstrated on the 4 000-tick workload); ASan (clang) run leak-free (exit 0, 600-tick safety shape; ≈1.64 ms/tick — 2.7× the canonical, the instrumentation the gate excludes); TSan `halt_on_error=1` race-free (exit 0) with one legitimate `system/budget_overrun` warn (BenchHash 2.04 ms vs its 2 ms budget under TSan's ≈4× slowdown — the engine's correct G-R5 observation, not a workload defect); full `ctest` green on all six local trees: 94/94 `build`/`build-clang`/`build-shared` (the 92 pre-existing + the two new gate entries), 93/93 `build-release` (one pre-existing Debug-only entry), 92/92 `build-asan`/`build-tsan` (the gate entries correctly absent there); zero new warnings under NFR-8.10; `tools/laige-include-lint` OK (47 source files, 1/10 deps), `tools/laige-determinism-lint` OK (28 sim files, 0 violations); CI: the PR's P0 jobs carry the gate entries in their full `ctest` (linux-gcc/clang default lane + the `ci:macos`/`ci:windows`-labeled windows-msvc/macos jobs + asan/tsan) — the M1-EXIT-01 gate records the final CI evidence (commit lands with the step's PR) | +| 2026-09-25 | M1-EXIT-01 | `—` / PR #52 | M1 exit gate: all 25 M1 steps confirmed against their Scope/Verify clauses — (1) sim-tick budget green on both SimMath backends: baseline report `docs/benchmarks/baselines/m1-sim-tick.md` (canonical Debug g++ 16.2.1: fpx16_16 mean 0.598405 / p99 0.606001 ms, fp32_pinned mean 0.591438 / p99 0.615179 ms — both PASS, 5.0×/8.1× inside the 3.0/5.0 ms targets) + CI (the CI reference machine, ubuntu-24.04, is the gate): PR-lane run 36152843429 (linux-gcc + linux-clang full ctest 94/94, including the two budget-gate entries `laige_bench_sim_tick_fpx16`/`laige_bench_sim_tick_fp32` — both budgets, both backends, exit 0) and run 36155154836 (macOS arm64 + Intel full ctest, the same two gate entries green on AppleClang); (2) replay bit-exact on CI across all P0 OS jobs and both backends: the determinism matrix `docs/benchmarks/determinism-matrix.md` (every P0 OS job asserts per-tick identity of the hello baseline on both backends — `hello_baseline_fpx`/`hello_baseline_fp32` — and the always-on detcheck job compares the cross-build pairs on both backends) — all-OS evidence: merge-lane run 36037080146 on master (2026-09-24: all 7 P0 jobs + detcheck green) and this PR's label-gated P0 matrix — Linux run 36152843429, macOS arm64/Intel run 36155154836, windows-msvc run 36156322585 — every job green, including the two hello-baseline tests in each OS job's full ctest and the detcheck job; (3) zero-alloc assertion green on the 10k workload: the `zero_alloc` ctest entry (M1-ALLOC-01) green in every P0 job's full ctest in the runs above (test logs in each run's job logs; the ASan/TSan lanes additionally archive the sanitizer reports) — and the M1-BENCH-01 measurement itself is zero-alloc-asserted: every tick of both backends (4 000 each — warm-up and measured) passed the engine's G-R1 per-tick allocation assertion in the debug trees (an allocating tick aborts the run). Deferred P1 items: none — every M1 step is P0. Progress Board: M1 25/25 complete (total 47/193); the board's stale Total row (41) corrected to 47 in the gate commit | ---