Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ The AI-relevant hardware the model runs on:
| `tests/` | Unit/integration tests, mirroring the `src/` module layout |
| `tools/` | Engine tools and CI scripts (fuzz runner, API manifest, determinism checker, lints) |
| `samples/` | Reference game projects (flagship isometric ARPG, platformer, lockstep arena, MMO demo zone) |
| `docs/` | Documentation; [decision index](docs/decisions/README.md) (full structure lands in M0-DOC-01) |
| `docs/` | Documentation; [index](docs/README.md) (full AGENTS §13 structure — getting started, concepts, API, guides, debugging, benchmarks, decisions, compatibility, testing) |

## Building

Expand Down Expand Up @@ -109,4 +109,6 @@ benchmark command `./build/bin/laige-bench --suite=<name>`). The test
- [AGENTS.md](AGENTS.md) — engineering contract (normative)
- [Roadmap](roadmap/README.md) — implementation checklist, progress board,
canonical commands
- [Architecture decisions](docs/decisions/README.md) — ADRs 0001–0003
- [Documentation index](docs/README.md) — getting started, concepts, API
contracts, guides, debugging, benchmarks, compatibility, testing
- [Architecture decisions](docs/decisions/README.md) — ADRs 0001–0004
102 changes: 73 additions & 29 deletions docs/README.md
Original file line number Diff line number Diff line change
@@ -1,16 +1,26 @@
# Laige documentation

Documentation index and navigation (DOC-001). The engine is at **M0**
(foundations): `laige-core` is the only populated module, and the
sections below mark what exists and what is still to land.
(foundations): `laige-core` is the only populated module. Every section
of the AGENTS §13 `docs/` tree exists; each entry below links what is
written and the "not yet written" section marks what is still to land.

## Build & tools
## Getting started

- [Building Laige](getting-started/building.md) — the source of truth for
the canonical build commands, build trees, options, compiler policy
(NFR-8.10), sanitizer builds (NFR-8.2), and the current M0 status.
Tool commands: `laige-fuzz`, `laige-bench`, `laige-detcheck`, the
`laige-api` manifest target, and the include-graph lint.
- [Building Laige](getting-started/building.md) — the source of truth
for the canonical build commands, build trees, options, compiler
policy (NFR-8.10), sanitizer builds (NFR-8.2), and the current M0
status. Tool commands: `laige-fuzz`, `laige-bench`, `laige-detcheck`,
the `laige-api` manifest target, and the include-graph lint.

## Concepts

- [Concepts index](concepts/README.md) — architecture, coordinates
(ARCH-008), lifecycle, threading, and determinism scope. **Not yet
written** (M0 is foundations only); the index names each planned
document and its interim home today (the `Vec2`/`Vec3` comments in
`src/laige-core/include/laige/sim_math.h`, ADR 0002, the per-API
contracts).

## API contracts (per public header)

Expand All @@ -28,47 +38,81 @@ sections below mark what exists and what is still to land.
- [Budget harness](api/budget_harness.md) — `Histogram`, `TimeIt`,
`budgetCheck`, the AGENTS §12 report format, and the `budgets.json`
schema (M0-CORE-08).
- [PRNG](api/prng.md) — `laige::Prng`: the splitmix64/LCG64 hybrid,
period, and determinism contract (M0-CORE-06).
- [PRNG](api/prng.md) — `laige::Prng`: xorshift128+ with splitmix64
seeding, substreams, period, and the determinism contract
(M0-CORE-06).
- [Determinism checker](api/detcheck.md) — the `laige-detcheck` tool and
the scenario hash-line contract (`<tick> <hash>` lines, two build
configurations) (M0-TOOL-02).

## Testing
## Guides

- [Testing conventions](testing.md) — test layout (module dirs mirror
`src/`, `<module>_tests` executables), the `regress_<short-id>`
regression-test convention, `laige-fuzz` target registration and CI
lane semantics, and the seed-handling convention for randomized tests
(M0-TEST-01).
- [Guides index](guides/README.md) — task-oriented usage and
optimization guides. **None yet**: they land with their milestones
(first headless project and determinism/replay in M1, profiling in
M1, rendering in M2, MMO server setup in M6/M7).

## Debugging

- [Debugging index](debugging/README.md) — what is usable today
(structured logging, the budget harness, `laige-detcheck`, sanitizer
builds, test seeds). The in-engine debug mode (AGENTS §15) lands with
the profiling work (M1-PROF-01, M2-PROF-01).

## Benchmarks

- [Benchmarks index](benchmarks/README.md) — methodology, baselines,
results, and the regression policy.
- [Benchmark methodology](benchmarks/methodology.md) — the AGENTS §12
report fields, the `budgets.json` field mapping, the baseline-file
convention, and the PRD §8.1 regression policy (M0-DOC-01).
- [Baselines](benchmarks/baselines/README.md) — recorded baseline
reports. **Empty so far** (all `budgets.json` entries have
`measured: 0`); the first, `m0-synthetic.md`, lands with M0-EXIT-01.

## Architecture decisions (ADRs)

- [ADR index](decisions/README.md) — 0001 (name and license), 0002
(deterministic math), 0003 (config JSON), 0004 (GoogleTest
vendoring).

## Compatibility

- [Compatibility](compatibility/README.md) — P0 platforms and
compilers (PRD §6; the CI matrix), the current machine-readable
formats (`budgets.json`, `laige-api.json`, `deps.lock`), and the
migration-guide status (none yet — pre-1.0, no breaking changes).

## Testing

- [Testing conventions](testing.md) — test layout (module dirs mirror
`src/`, `<module>_tests` executables), the `regress_<short-id>`
regression-test convention, `laige-fuzz` target registration and CI
lane semantics, and the seed-handling convention for randomized tests
(M0-TEST-01).

## Not yet written (honest status)

- `concepts/` — architecture, coordinates (ARCH-008; lands as
`docs/concepts/coordinates.md` with M0-DOC-02 — until then the
coordinate system is documented in the `Vec2`/`Vec3` comments of
`src/laige-core/include/laige/sim_math.h`), lifecycle, threading.
- `guides/` — task-oriented usage (first game, profiling, determinism).
- `debugging/` — debug mode (AGENTS §15 lands in M3), logging in
production, troubleshooting.
- `benchmarks/` — method, baselines, and the regression policy
(the harness exists — `laige-bench`, M0-CORE-08 — but the recorded
baselines land with M1 workloads).
- `compatibility/` — platform/compilers/formats matrix (the P0 matrix
is in [building.md](getting-started/building.md) for now).
- `concepts/` — the architecture, coordinates, lifecycle, threading, and
determinism concept documents (the [index](concepts/README.md) names
each and its interim home).
- `guides/` — task-oriented usage (first game, profiling, determinism)
— see the [index](guides/README.md).
- `debugging/` — the in-engine debug mode (AGENTS §15; profiling
foundations land in M1, render observability in M2).
- `benchmarks/baselines/` — no measured baselines yet; the first lands
with M0-EXIT-01.
- `compatibility/` — no persistent data formats or migration guides yet
(they land with M1 replay and M6/M7 networking).
- Per-module API docs for the M1+ modules (`laige-sim`, `laige-render`,
`laige-assets`, `laige-net`, `laige-server`, `laige-script`,
`laige-editor`) — they land with their modules.

## Related

- [Roadmap index](../roadmap/README.md) — the M0/M1/... step plan;
- [Roadmap index](../roadmap/README.md) — the M0/M1/… step plan,
progress board, and change log;
[M0 foundations](../roadmap/M0-foundations.md) is the current
milestone.
- `AGENTS.md` — the engineering contract this documentation implements.
- `PRD.md` — the product requirements.
22 changes: 22 additions & 0 deletions docs/benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Benchmarks

Benchmark methodology, baselines, results, and the regression policy
(AGENTS §13). Performance claims are reproducible measurements
(CORE-001); this section is where they are recorded.

- [methodology.md](methodology.md) — the **normative methodology**
(M0-DOC-01): the AGENTS §12 report fields, how `budgets.json` entries
map onto them, the baseline-file convention, the PRD §8.1 regression
policy, and workload discipline.
- [baselines/](baselines/README.md) — recorded baseline reports
(`<milestone>-<workload>.md`). **Currently empty**: every
`budgets.json` entry has `measured: 0`. The first baseline,
`baselines/m0-synthetic.md`, lands with M0-EXIT-01, and
`baselines/m1-profiler-cost.md` with M1-PROF-01.
- Per-milestone performance results (M1 10k-entity tick, M2 50k-sprite
scene, M6/M7 zone server, …) land here as their milestones close —
see the [roadmap progress board](../../roadmap/README.md).

The measurement tool is `laige-bench` (M0-CORE-08): API contract in
[api/budget_harness.md](../api/budget_harness.md), canonical command in
[building.md](../getting-started/building.md).
20 changes: 20 additions & 0 deletions docs/benchmarks/baselines/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Baselines

Recorded baseline reports — the durable before/after history of the
`budgets.json` budgets (convention:
[../methodology.md](../methodology.md) §4). One file per recorded
measurement:

```text
<milestone>-<workload>.md e.g. m0-synthetic.md (M0-EXIT-01)
```

Each baseline file contains the full AGENTS §12 metadata (hardware, OS,
compiler + version, build type, flags, workload, warm-up, sample count,
summary statistics), the stable 4-line `budgetCheck` report verbatim,
and the exact command + commit that produced it. Baseline files are
**immutable**: superseding a baseline adds a new file and updates
`measured` in `budgets.json` — it never edits an existing baseline.

No baselines are recorded yet (every `budgets.json` entry has
`measured: 0`); the first lands with M0-EXIT-01.
159 changes: 159 additions & 0 deletions docs/benchmarks/methodology.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,159 @@
# Benchmark methodology

The normative measurement methodology for Laige performance work
(M0-DOC-01; AGENTS §12 report requirements, CORE-001 "measure first",
PRD §8.1 budget policy). It defines what every performance report must
contain, how the `budgets.json` entries map onto it, the baseline-file
convention, and the regression policy that makes budgets CI-enforced.

The measurement machinery is the budget harness
([api/budget_harness.md](../api/budget_harness.md)): `laige::Histogram`,
`laige::TimeIt`, `loadBudgets`, `budgetCheck`, and the operator tool
`laige-bench` (canonical command in
[building.md](../getting-started/building.md)).

## 1. Required report fields (AGENTS §12)

Every performance report — a baseline file under
[`baselines/`](baselines/README.md), a budget result archived with a
milestone, or a before/after pair for an architectural change
(TEST-009) — MUST record:

| # | Field | What it records |
|---|---|---|
| 1 | Hardware | CPU model/class, RAM, GPU where relevant |
| 2 | OS | Name and version, architecture (x64 / arm64) |
| 3 | Compiler and version | e.g. `g++ 16.2.1`, `MSVC 2022` |
| 4 | Build type | `Debug` (canonical) or `Release`; sanitizer trees named explicitly (`build-asan`, `build-tsan`) |
| 5 | Relevant flags | The engine policy (NFR-8.10) always applies; record any additional flags (e.g. the SimMath pinned set, a sanitizer flag set) |
| 6 | Dataset / workload | The named workload and its defining parameters — the `workload` text of the budget entry is the reference |
| 7 | Warm-up | Iterations discarded before sampling (`--warmup`) |
| 8 | Sample count | `n` — the number of recorded samples (the `n=` of the stats line) |
| 9 | Summary statistics | min / mean / p50 / p95 / p99 / max over the stored window (the `stats:` line) |
| 10 | Before / after | The checked budget's `before=` (last recorded — the entry's `measured`) and `after=` (this run) values |

Reporting rules (AGENTS §12, PERF-009, PERF-010):

- **Percentiles, not just averages.** Tail latency is a first-class
result: report p95/p99 and max alongside the mean, never the mean
alone.
- **Prefer repeatable automated benchmarks over ad hoc timings.** Runs
go through the canonical `laige-bench` command so a later reader can
reproduce the number byte-for-byte.
- **Representative workloads** — small, medium, and stress — per
PERF-010; a microbenchmark alone cannot validate an architectural
change.
- **Never change a benchmark solely to make a regression disappear.**
If the workload no longer represents the product, revise it in a
documented step (and the PRD §8.1 table with it, §5 below).

## 2. How `budgets.json` maps to the report

`budgets.json` (repo root) is the machine-readable budget table; schema
version 1, strictly validated by `loadBudgets` (ARCH-007; the schema is
documented in [api/budget_harness.md](../api/budget_harness.md)). Every
entry carries exactly six fields, and each maps to a report role:

| `budgets.json` field | Report role |
|---|---|
| `name` | The budget's identity — the `budget=<name>` of the report's first line and the section/baseline name |
| `metric` | Which statistic the check evaluates (`mean` / `min` / `max` / `p50` / `p95` / `p99`) — the report's `metric=` field |
| `unit` | The unit of `target`, `measured`, and every measured value |
| `target` | The hard PRD §8.1 limit — an at-most upper bound; `target == 0` is a **hard-zero budget** (e.g. `sim_heap_allocs`), not "not set" |
| `measured` | The last recorded value — the **before** number of the before/after pair. Updated when a run is recorded as the new baseline; `0` is the M0 convention for *not yet measured* |
| `workload` | The workload the budget applies to — the reference text of report field 6. The caller harness measures exactly this workload and echoes it in `context: workload=` |

`budgetCheck` produces the before/after pair automatically: `after` is
the current run's value of `metric` over the histogram window, `before`
is the entry's `measured` field. Recording a run as the new baseline
means writing its `after` value back into `measured` in `budgets.json`
and committing the baseline file (§4) and the table in the **same
change** (CORE-006).

## 3. The harness report and §12 field coverage

`budgetCheck` emits the stable 4-line report (LOG-001 machine-greppable;
the format's single source of truth is `laige::formatStatsLine` — see
[api/budget_harness.md](../api/budget_harness.md)):

```text
budget=<name> result=<PASS|FAIL|NO_SAMPLES> metric=<m> unit=<u>
after=<v> before=<v> target=<v>
stats: n=<n> min=<v> mean=<v> p50=<v> p95=<v> p99=<v> max=<v>
context: workload=<s> build=<s> machine=<s> warmup=<n>
```

Coverage of the §1 field table:

| Report field | Provided by |
|---|---|
| 10 (before / after / target) | the harness, from the entry and the histogram |
| 8 (`n`), 9 (statistics) | the `stats:` line |
| 6 (workload), 7 (warm-up) | the `context:` line — **caller-supplied**: the operator sets `--workload`/`--warmup` |
| 4 (build), 1 (machine, partially) | the `context:` line — the tool supplies compiler + build type; the operator supplies the machine (`LAIGE_BENCH_MACHINE`) |
| 2 (OS), 3 (compiler version), 5 (flags) | **not in the report** — a baseline file or archived CI log must add them |

Consequently a baseline file is: the 4-line report verbatim, plus the
remaining §12 fields (hardware, OS, compiler + version, flags), plus the
exact command and the commit it was measured on.

## 4. Baseline files

- **Location and naming:** `docs/benchmarks/baselines/<milestone>-<workload>.md`
(e.g. `m0-synthetic.md` — written by M0-EXIT-01 — and
`m1-profiler-cost.md`, written by M1-PROF-01).
- **Content:** the complete §1 record for one measurement: the §12
fields, the stable 4-line report verbatim, the exact canonical command,
and the measured commit.
- **Immutability:** baseline files are records, not live state. A new run
supersedes an old baseline only by adding a new baseline file and
updating `measured` in `budgets.json` — never by editing an existing
baseline (the before/after history must survive for regressions and
audits).
- **`budgets.json`'s `measured`** always holds the *latest* recorded value
of each budget (its `0` entries mean "not yet measured" until the
subsystem that owns them lands).

## 5. Regression policy (PRD §8.1, NFR-8.1)

- **Budgets are part of CI.** A PR that regresses any budget by **> 10%**
against its last recorded value (**`before`**) — or breaches the
absolute **`target`** — fails CI unless the budget is revised via a
PRD revision.
- **CI-gateable:** `laige-bench --budget=<name>` exits `0` on pass and
**`2` on a failed budget check** (`1` = usage or load failure); the
CI lane asserts the exit code.
- **The 10% band needs a baseline:** it applies only when
`before > 0`. A first measurement (`before == 0`, not yet measured)
must still pass the absolute target.
- **Hard-zero budgets have no band:** `target == 0` passes only when the
measured value is exactly 0 (any positive measurement is a failure).
- **`NO_SAMPLES` is a failure, not a skip:** a workload that recorded
nothing is a broken harness (CORE-008 — never silent).
- **Accepted budget revisions** change `budgets.json` in the same PR as
the PRD revision, and are noted in the milestone change log
(DOC-007) — a budget number is never silently moved.

## 6. Workload discipline

- **Deterministic workloads first.** The M0 reference workload
(`laige-bench --suite=synthetic`) is deterministic by construction
(fixed Marsaglia LCG64 constants, no RNG, no allocation) — which is
what makes its baseline reproducible across machines and commits.
- **No silent window truncation.** Set the histogram
`capacity >= runs` for a claim over "all samples"; if a rolling window
is used deliberately, check `totalRecorded()` and say so in the report
(the harness never hides a drop — CORE-008).
- **Warm up, then sample.** Discard the first N iterations (canonical
`--warmup=100`) so first-touch costs do not pollute samples; record N
in the report.
- **Tail latency always.** Budgets are at-most upper bounds on a named
percentile (or max/mean) — the `metric` field of the entry — so the
reported statistic must be the one the budget checks, not a friendlier
one.
- **Randomized workloads use the repo-wide seed convention**
([testing.md §4](../testing.md)): fixed default seed `0x1F055EED`,
overridable, so CI runs of the same commit are comparable.
- **Record the environment.** Runs on different P0 platforms are not
comparable numbers: the §12 hardware/OS fields exist so a baseline
never travels without its machine.
Loading
Loading