Skip to content

Fix macOS/Windows CI: platform-independent churn cost check; MSVC C4100/C4702 - #20

Merged
offdev merged 1 commit into
masterfrom
fix/ci-macos-windows
Sep 14, 2026
Merged

offdev merged 1 commit into
masterfrom
fix/ci-macos-windows

Conversation

@offdev

@offdev offdev commented Sep 13, 2026

Copy link
Copy Markdown
Owner

Root cause (two independent failures, both red since M1-ECS-03)

macOS (macos-14 + macos-15): ArchetypeChurn.TenKEntitiesAddRemoveChurnZeroAllocAndFlatCost fails the wall-clock assertion p99 < 3 x p50. The shared macOS runners' wall-clock tail alone reaches ~5x the median — measured 2026-09-13: p99/p50 = 4.9 on both macOS runners (max up to 2.25 ms) vs 1.98 on Linux — while the work done is identical on every platform. A raw time-ratio assertion does not travel across P0 platforms; docs/benchmarks/methodology.md §6: "runs on different P0 platforms are not comparable numbers."

Windows (MSVC, /W4 /WX): two warning classes, all in the laige-sim test TUs:

  • C4100 unreferenced parameter — World::each(F&& fn, Acc... accesses): the access-tag values were never referenced (only their types are used, via static_asserts and template arguments).
  • C4702 unreachable code (833 instances) — the if constexpr (Lo < Hi) recursion helpers in archetype_tests.cpp (×3) and component_registry_tests.cpp (×1): in the Lo < Hi instantiations the trailing return true; is unreachable.

Fix

macOS — the cost check is now platform-independent (the zero-overflow / zero-growth / zero-allocation assertions are unchanged):

  1. Deterministic work KAT. New ArchetypeStats::totalRowShifts counter (rows moved by attachSlot/removeRow tail shifts). The window's total is a seed-invariant constant of the workload: each op moving slot s shifts exactly 2 × (9999 − s) rows (with all slots live and partitioned between the two archetypes, s's rank in each archetype sums to s), so the window shifts exactly 10000 × 9999 = 99,990,000 rows for any seed. Verified against an independent cost-model simulation: 99,990,000 for every seed swept. A regression that adds or drops O(rows) work in the move path (hidden scan, swap-remove rewrite, double copy) breaks this count; no seed can.
  2. Per-row wall-clock floor (the "measured, not assumed" half): window time per shifted row ≤ 200 ns/row. Measured: ~24 ns/row Linux x64 Debug (g++/clang++ -O0), ~7 ns/row macOS Debug (AppleClang -O0), 60.4 ns/row under ASan (slowest configuration observed). VM preemption spreads across the whole window (100 ms of deschedule over ~1e8 rows is < 1 ns/row), so the floor is immune to the shared-runner noise that inflates raw percentiles, while still catching an order-of-magnitude per-row regression the row count cannot see.

The machine-greppable stats line is kept, plus a new archetype-churn work: rows_shifted=… ns_per_row=… line. docs/api/archetype.md updated (DOC-003); laige-api.json regenerated (api-real-tree green).

Windows:

  • World::each access-tag pack is now unnamed in declaration and definition (C4100 fires only for named unused parameters; semantics unchanged — callers pass the tags positionally as before).
  • The four recursion helpers get an explicit else branch: a discarded if constexpr branch is never emitted, so there is no unreachable trailing statement in either instantiation (C4702).

Validation

  • Local (sandbox): full ctest green on all four trees — build (g++ Debug), build-clang (clang++ Debug), build-asan (clang++ Debug ASan+UBSan, alloc-counter path), build-release (Release) — 37/37 each. laige-api-scanner --check OK.
  • Churn evidence per tree (rows_shifted = 99990000 in all): g++ Debug 24.2 ns/row · clang++ Debug 24.2 ns/row · ASan 60.4 ns/row · Release 0.4 ns/row.
  • macOS/Windows cannot be built in this sandbox; the fixes target the exact compiler diagnostics from the failing job logs, and the macOS assertion no longer depends on wall-clock percentiles at all. The full P0 matrix will confirm on the merge run.

…ck; MSVC C4100/C4702

macOS (macos-14/macos-15) and Windows (windows-2022) jobs red since
M1-ECS-03 landed:

macOS: ArchetypeChurn failed the wall-clock assertion p99 < 3x p50.
The shared macOS runners' wall-clock tail alone reaches ~5x the
median (measured: p99/p50 = 4.9 on both macOS runners vs 1.98 on
Linux), while the work done is identical on every platform — a raw
time-ratio assertion does not travel across P0 platforms
(methodology.md section 6: 'runs on different P0 platforms are not
comparable numbers'). The assertion is replaced by two
platform-independent checks of the same property:
  (1) a deterministic work KAT: the window's total row-shift count
      is seed-invariant (each op moving slot s shifts exactly
      2 x (9999 - s) rows, since s's rank in each archetype sums to
      s), so the window shifts exactly 10000 x 9999 = 99,990,000
      rows for ANY seed; verified against a cost-model simulation
      (99,990,000 for every seed swept). This needs the new
      ArchetypeStats::totalRowShifts counter (attachSlot/removeRow
      tail shifts).
  (2) a per-row wall-clock floor (<= 200 ns/row): measured
      ~24 ns/row Linux x64 Debug, ~7 ns/row macOS Debug, 60 ns/row
      under ASan; VM preemption spreads across the whole window, so
      the floor is immune to the shared-runner noise while still
      catching order-of-magnitude per-row regressions the row count
      cannot see.
The machine-greppable stats lines are kept (a 'work' line added).
docs/api/archetype.md updated (DOC-003); laige-api.json regenerated.

Windows: two warning classes fatal under /WX:
  - C4100 unreferenced parameter: World::each's access-tag pack
    values were never referenced (only their types are used) — the
    pack is now unnamed in the declaration and definition.
  - C4702 unreachable code: the if constexpr (Lo < Hi) recursion
    helpers in archetype_tests.cpp and component_registry_tests.cpp
    had a trailing return unreachable in the Lo < Hi
    instantiations — explicit else branches (discarded statements
    are never emitted).
@offdev
offdev merged commit 155fbff into master Sep 14, 2026
10 checks passed
offdev added a commit that referenced this pull request Sep 14, 2026
…tform-dependent lastMs assertion (#28)

The windows-msvc, macos-intel, and macos-arm64 jobs are red on the
M1-SYS-03 merge (run 82). Two independent, platform-sensitive
defects, both in tests/laige-sim/system_timing_tests.cpp:

macOS Intel (macos-14, AppleClang 15) — build error:
  system_timing_tests.cpp:125: error: compound assignment to object
  of volatile-qualified type is deprecated
  [-Werror,-Wdeprecated-volatile]
The P1152R2 volatile deprecation (compound assignment to a
volatile-qualified type) is diagnosed by AppleClang 15 and fatal
under the engine policy's -Werror, while the other P0 job
toolchains do not diagnose it (Clang 18 on ubuntu-24.04,
AppleClang 16 on macos-15, MSVC 2022) — so only macos-intel failed
to build. The burn-loop sink is now updated with an ordinary
assignment (gBurnSink = gBurnSink + 1;), the pattern already
established in budget_harness_tests.cpp's Spin() helper. The volatile
read-modify-write remains, so the loop still defeats elimination.

Windows (windows-2022) + macOS arm64 (macos-15) — test failure:
  SystemTiming.HealthyTicksLogNothingAndTrackStats:
  EXPECT_GT(st.value().lastMs, 0.0) — actual 0 vs 0
A noop system runs in tens of nanoseconds; on platforms whose
steady_clock tick is coarser than that (the start and end reads land
on the same tick) the measured run time is exactly 0.0 ms — a
legitimate sub-resolution reading (wall-clock resolution is
platform-sensitive; ARCH-009), not a failure state. The assertion
was process-state-dependent: the unfiltered laige-sim_tests entry
failed on Windows while the filtered system_timing entry passed, and
vice versa on arm64 — the same flake class as the pre-M1-ECS-05
churn cost checks fixed in #20. The lower bound becomes
EXPECT_GE(st.value().lastMs, 0.0); the tracked-state property is
already pinned by the runs counter, and the upper bound (< 1.0 ms,
under the 1 ms budget) is unchanged.

Docs (DOC-006/007): system.h (SystemTimingStats::lastMs and
SystemTimingRecord::lastMs) and docs/api/system_timing.md now state
that a run shorter than the platform's steady_clock tick measures
as exactly 0.0 ms and must not be treated as "not measured".

laige-api.json: regenerated (manifest entries carry declaration line
numbers; the system.h comment additions shifted four declarations).
The api-manifest job's check passes against the regenerated
manifest.

Verification (local, g++ 16.2 / clang 22.1, Debug trees):
  - full ctest suite: 43/43 green on both toolchains
  - include-lint: OK
  - laige-api-scanner --check: up to date
The Windows and macOS fixes are verified by this PR's ci:windows
job, and (after the label swaps to ci:macos) the two macOS jobs.
offdev added a commit that referenced this pull request Sep 15, 2026
…SecondRunFailsWithoutLogging (#33)

* [M1-HEAD-01] Fix macOS CI: platform-dependent zero-warn assertion in SecondRunFailsWithoutLogging

The macos-arm64 and macos-intel jobs (merge matrix) are red on the
M1-HEAD-01 merge:
  [  FAILED  ] EngineRun.SecondRunFailsWithoutLogging
  engine_tests.cpp:411: Expected equality of these values:
    sink->entries.size() (Which is: 1) vs 0u

The test asserted zero Warn-or-above entries over the WHOLE capture
window - the first wall-clock-paced run_headless(1, 1) at 60 Hz (one
16.7 ms tick) plus the stopped-state second run. On a loaded runner a
frame can run longer than one tick of simulation time, and the loop
then warns loop/tick_dropped - the documented M1-LOOP-01 overload
behavior (the BoundedRun tests in the same suite deliberately do not
assert the drop count for this reason). The macOS runners' wall-clock
tail (measured in the same merge run: dropped=3 in BoundedRun,
dropped=2 in DoubleShutdownAfterRun) makes the warn near-certain
there, while the faster Linux runners happened to stay under the
16.7 ms frame - a wall-clock-dependent assertion that does not travel
across P0 platforms (the #20/#26/#28 CI-fix precedent).

The test's actual property is only about the SECOND run: a stopped-
state failure is a pure failure, no log (the GameLoop's moved-out
precedent). The assertion is now scoped to that run - the sink entry
count is recorded after the first run, and the second run must add
none. The first run's legitimate tick_dropped warn (its own
accounting, rate-limited per LOG-004) no longer taints the
stopped-state check.

No engine/loop code change: the second run still returns
InvalidArgument before any log call (engine.cpp run_headless,
stopped-state branch). Verified locally on the gcc and clang trees
(47/47 ctest, the EngineRun suite green).

Verified by this PR's ci:macos job (macos-arm64 + macos-intel).

* Retrigger CI: ci:macos label applied to the PR

The first pull_request event fired when the branch was pushed, before
the PR existed and before the ci:macos label was attached, so that
run's macOS jobs were skipped by the label selector. This empty commit
re-fires the pull_request event with the label in place (the
concurrency group cancels the label-less run).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant