An initiative-scoped workflow system for coding agents: it takes a major idea from intake through architecture, spec, plan, plan regression, execution, review, and verification — with a strict per-initiative ID namespace, separated thinking and execution stores, a script-enforced contract at every state transition, and evidence-backed completion.
Nothing floats. Every document, task, review, verdict, and ruling carries an ID that names the initiative it belongs to. A body of work gets one Initiative; the initiative owns a folder, an ID namespace, and every document produced about it.
Nothing is enforced by hope. Where a workflow could silently drift — a run row lying about the ledger, a task completed without review, a plan naming paths only scripts may resolve — a script checks it and exits non-zero. Contracts enforced by prose are hopes; contracts enforced by scripts are contracts.
Works with any agent that can read skill files from a directory — Claude Code, Codex, omp, or anything similar. The skills are markdown; the scripts are POSIX bash.
One line, any of 20+ agents (Claude Code, Codex, Cursor, Copilot, Gemini
CLI, OpenCode, Amp, Goose, and more) — skills
discovers every skill in the repo and wires them into your harness:
npx skills add Atri10/executorOr clone and copy skills/ into whatever directory your agent loads from:
git clone git@github.com:Atri10/executor.git
cp -R executor/skills/* <your-agents-skills-dir>/Pin to a release tag instead of main for stability:
git clone --branch v0.7.1 git@github.com:Atri10/executor.gitPaste this into your agent — Claude Code, Cursor, Aider, Codex, omp, or any other harness. It discovers the right skills directory itself and verifies the install:
Install The Executor skill library for me:
1. Clone https://github.com/Atri10/executor.git into a temp directory
(use --branch v0.7.1 for the latest release, or default branch for main).
2. Find my agent's skills directory. Candidates, in order — use the first
that exists, or ask me if none do:
- ~/.omp/agent/skills/ (omp)
- ~/.claude/skills/ (Claude Code)
- ~/.cursor/skills/ (Cursor)
- .claude/skills/ (repo-local Claude Code)
- ~/.aider/skills/ or as my harness documents
3. Copy every directory from the clone's skills/ folder into that skills
directory (each is one skill: skills/executor, skills/executor-spec, ...).
4. Verify: run bash <skills-dir>/executor/scripts/exec with no arguments
inside the clone — it must print the resume digest (or a clean "nothing
in flight"), proving the script tree resolves. Then run
bash <skills-dir>/executor/scripts/exec-graph check inside the clone —
it must report 0 failed, proving the whole script tree is intact. Also
confirm every copied skills/<name>/ directory contains a SKILL.md (count
the directories, not a remembered number — the skill set grows, and a
stale count here fails an install that is actually fine).
5. Tell me which directory you installed into, and how to invoke the
router in my harness (usually /skill:executor or just asking for
"the executor").
Do not modify any file inside the clone or the skills directory other
than the copy operation itself.
Then invoke the root router:
/skill:executor
or just say "start an initiative" — normal requests route by each skill's frontmatter description.
| Skill | Phase | Output |
|---|---|---|
executor |
Router + contract | Loads the right phase skill, defines the ID namespace |
executor-initiative |
Intake | Initiative folder, charter, registry entry, initiative branch |
executor-discovery |
Discovery | Research notes, options comparison |
executor-architecture |
Architecture, Design | Architecture, ADRs, interfaces, component designs |
executor-spec |
Specification | Spec, risks, verification strategy (one row per requirement) |
executor-planning |
Planning | Plans with tasks, linted before the gate |
executor-plan-regression |
Plan regression | Plan-set audit, repairs, clearance summary — gates execution |
executor-execution |
Execution | Task dispatch, ledger, reports, run registry |
executor-review |
Review | Per-task and whole-branch verdicts, findings, fix loops |
executor-verification |
Verification | Evidence-backed proof each requirement holds |
executor-handoff |
Handoff | Human decision menu: merge, PR, or keep the branch |
executor-brainstorm |
Cross-phase | Recorded divergent ideation sessions; visual companion optional |
executor-critique |
Every phase (gate) | Per-component audit, repair, re-audit; gates the phase that produced the artifacts |
Phases compress, they never vanish. A small initiative can produce a charter and a spec in one exchange and skip discovery — but skipping is a stated decision recorded in the charter, not an omission.
The controller runs one loop and holds no judgment: it runs exec, reads
the single action word it prints, does that, and goes back. Every arrow
below is either a script decision or a dispatched subagent. There is no
step where the controller does the work.
flowchart TD
start(["exec"]) --> enter{"PHASE-ENTER phase"}
enter --> author["exec-dispatch --role author"]
author --> artifact["artifact on disk"]
artifact --> check["exec-critique check"]
check -->|"not clear"| loop["AUDIT then REPAIR then re-AUDIT"]
loop --> check
check -->|"clean or waived"| gate{"PHASE-GATE phase"}
gate -->|"human gates it"| next["next phase"]
gate -->|"exec-present gate card"| next
gate -->|"autonomous declared"| auto["exec-gate --auto"]
auto -->|"refused"| gate
auto -->|"auto-passed"| next
next --> start
start -->|"every phase resolved"| done(["DONE"])
The phase order, and which of them carry a critique gate. Every authoring phase does; the two critique phases audit the others.
flowchart LR
I["intake"] --> D["discovery"]
D --> A["architecture"]
A --> G["design"]
G --> S["specification"]
S --> P["planning"]
P --> PR["plan-regression"]
PR --> E["execution"]
E --> R["review"]
R --> V["verification"]
V --> H["handoff"]
classDef gated fill:#1f3a5f,stroke:#4a90d9,color:#fff
classDef critic fill:#3d2f1f,stroke:#d9a04a,color:#fff
class I,D,A,G,S,P,E,V,H gated
class PR,R critic
Each plan's tasks run as a loop, and a task result reaches run state only through a gate. A worker that stops is escalated through a fixed ladder before anyone is replaced.
flowchart TD
step(["exec step PLAN_FILE"]) --> act{"action"}
act -->|"RUN-START"| start["exec-run PLAN start"]
act -->|"DISPATCH"| disp["exec-dispatch PLAN --role impl"]
disp --> impl["spawn IMPL on the printed PROMPT"]
impl --> worker[("worker runs on its own branch")]
worker -->|"report lands"| rep{"verdict?"}
rep -->|"none, no reviewer"| rev["REVIEW: exec-dispatch --role review"]
rev --> reviewer[("reviewer writes verdict")]
reviewer --> rep
rep -->|"none, reviewer in flight"| waitrev[("WAIT — no livelock")]
rep -->|"unclean"| fix["FIX: exec-dispatch --role fix"]
fix --> worker
rep -->|"clean"| commit["REPORT: exec-report gates and commits"]
commit --> next{"more tasks?"}
next -->|"yes"| step
next -->|"no"| gate["GATE-STAGE: exec-run complete"]
act -->|"REVIVE / REDISPATCH"| ladder["exec-ladder rewrites the rung"]
ladder --> worker
act -->|"ADJUDICATE"| adj["exec-adjudicate builds the evidence pack"]
adj --> sup["spawn SUPERVISOR on the printed PROMPT"]
sup --> step
act -->|"WAIT"| idle[("idle one turn — the emit names the wake")]
act -->|"DONE"| fin(["plan complete"])
A worker's status reaches run state through two gates, never through the
controller's memory: the report's result: frontmatter (blocked routes to
ADJUDICATE, needs-context to ASK) and exec-report's verdict gate. A
report with no verdict yet emits REVIEW — and WAIT while that reviewer
runs, which is what stops a healthy run from reading as a stalled one.
exec-briefextracts one task's text into a self-contained brief — the implementer reads requirements in one call, and task text never passes through the controller's context.exec-contextassembles everything the brief cannot know: the exact signatures earlier tasks provide, the current surface of the files being modified, the binding global constraints, and the rulings that touch the task's files. Implementers start working without exploring.- Two-verdict reviews (spec compliance + code quality) from a reviewer who never trusted the implementer's report, writing a verdict file — not a chat message that vanishes on the next summarization.
- Non-code tasks still get reviewed: a docs-only or evidence-capture task is judged on its report vs its brief, with the same mandatory verdict file.
Findings are severity-graded with worked calibration examples, fixed in rounds (1–3 resume the original implementer; 4–5 escalate to a fresh, more-capable model), re-reviewed scoped to the fix diff, and at the cap adjudicated by recorded ruling — never silently dropped. Re-reviews check the fix addressed the root cause, and whether any test was weakened.
The implementer contract asks for the strongest feasible evidence for every behavior change: a watched failing test where a test harness exists, a named alternative instrument where it does not (CLI fixture run, parse/render check, exercised UI), and an explicit NOT-RUN/UNAVAILABLE record where nothing feasible exists. Reviewers verify evidence, and a reviewer who cannot name the failure a demanded test would catch does not get to demand it.
Per-plan review reads one file. The defects that actually ship live
between files: an Assumes section promising a signature the
predecessor plan never produces, a spec requirement no plan's Covers:
claims, two plans provisioning the same queue with different shapes.
executor-plan-regression runs between planning and execution and reads
the whole set against the spec, the interface contracts, and each other —
coverage, Assumes/Produces closure, ordering, constraint propagation,
file-map collisions, vocabulary. Findings are repaired in the plan files
themselves, then re-audited; the clearance summary lives at
.executor/<INIT>/plan-regression/summary.md. Execution cannot start
until every plan is clean or the human has explicitly waived a
finding — enforced by the phase-order machine, by exec-run start, and
by exec-store-check.
Every role that dispatches a subagent — implementer, task reviewer,
re-reviewer, final reviewer, plan auditor, plan repairer, plan
re-auditor, evidence runner, prior-art scout, concept explorer, design
critic — has a registered prompt template with its identity grammar, a
required model: line, and a filled [PLACEHOLDER] contract. An agent
briefed from the dispatcher's memory is a dispatch defect:
validate-skills.sh fails on a dispatch role with no registered template,
a template missing its ## Self-Critique Before You Return /
## Verification / ## What You Return sections, or a prompt file nobody
registered. The dispatch registry lives in
references/layout.md.
Each phase skill carries a ## Self-Critique section — an adversarial
pass over the artifact it just wrote — and a ## Verification section —
the commands that prove it, run in-session with output cited at the gate.
A gate claimed before both ran is not claimed. The same two sections are
required inside every dispatch template, so a subagent checks its own
output before returning a status line the controller will trust.
The pipeline has eleven phases. Before this, only two of them could
reject a bad deliverable — plan-regression and review. design and
verification had no artifact gate at all, and architecture and spec ran
adversarial self-critiques whose own skills refused the word critic,
correctly: one agent grading its own work is not review.
executor-critique generalises the plan-regression contract to all nine
authoring components. Each phase's artifacts are audited as a set — the
defects worth finding live between documents — and the phase cannot be marked
passed until that audit is clean or the human has waived it in writing.
The gate is a script, and it is harder than the stage it generalises from:
verdict: PASSbeside a non-zerohigh/mediumis refused, not believed. PASS means zero HIGH and zero MEDIUM, and the check recomputes that from the auditor's own counts.- The three-round cap is enforced by the script.
audit 04exits 2 with a message. A cap nobody counts is a suggestion, and a suggestion is how a fourth round appears with nobody able to say when the loop started. - A waiver needs both a note in the row and an initiative ruling naming
the artifact. A bare
waivedcell is indistinguishable from a controller self-granting one. - An empty clearance is not a cleared one — for document components and run-axis components alike.
The controller never clears its own work. exec-initiative runs the check on
the controller's behalf and refuses the phase, and the refusal names what to
do about it rather than leaving a way around it.
No reviewer prompt in the system ever opened the architecture store. Its
inputs were the spec, the plan, the ledger, reports, verdicts and the diff —
so an implementation satisfying every R-nn and every C-nn while violating
the ARCH passed every gate in the system. Adapters importing each other, an
ORM row crossing into policy, a seam renamed from its IFCE: there was no point
at which anything could see it.
Two seats now catch it. The code component's check catalog makes the
architecture, IFCE and design stores required inputs of the implementation
critique, and the final reviewer's prompt takes them as required inputs with an
architecture-conformance check — where the spec and the architecture disagree,
that is a finding, and the diff alone cannot say which side is wrong.
An initiative can declare which phase gates a human agreed to let go unattended:
| Phase | Mode | Why |
|---|---|---|
| intake | deny | the charter is the human's to approve |
| execution | allow | every stage is gated; workers are dispatched, not self-approved |exec-gate INIT PHASE --auto clears a gate only when the policy permits it,
and every ambiguity resolves to deny: no policy file, enabled: false, an
unlisted phase, or a mode spelled wrong. A typo in a table cell must not become
an unattended phase gate. A pick-class phase (design, specification) in
allow mode still requires a scored verdict already on disk — a script can
count files, it cannot choose between designs, so it declines rather than
guesses.
An auto-pass is recorded as **auto-passed** <date>, visibly distinct from a
human passed, and it must clear the same artifact gate a human pass does. A
human who disagrees overturns it with one command, which reopens the phase:
exec-initiative phase INIT-0004 superseded specification "the design was wrong"Inside a phase the workflow does not wait on a human — but it does not
silently decide everything either. Mechanical or reversible calls are
ruled and logged (exec-ruling, mirrored into .local/decisions/).
Decision-class calls — contract conflicts, scope cuts, irreversible
choices, product judgment — are asked on the spot through the harness's
question affordance, blocking only the affected lane, with the question
recorded beside the ruling (--answered). Only four things stop a run:
an irreversible operation, a security-sensitive action, a side effect
outside the worktree, or a defect where every path forward is a guess.
| Script | Owns |
|---|---|
exec |
The single entry point: bare exec prints the resume digest, exec <name> dispatches to exec-<name>, exec verbs prints the decision table, exec roles the dispatch registry. The controller never needs the scripts directory path |
exec-initiative |
Allocate initiative IDs, scaffold folders, phase log, initiative branch (branch INIT-0004); phase … check validates a gate without writing it |
exec-id |
Next free ID of any type — allocation never guesses |
exec-plan-lint |
Planning gate: rejects literal store paths in plans, task headings without IDs, missing or empty spec/interfaces/tasks/execution_mode, task-count mismatch, and over-specified task bodies (impl-language fences >40 lines or >60% of a body) |
exec-workspace |
Resolve and seed a plan's execution workspace: ledger, rulings, preflight scan, dispatch log; seeds the registry row ready |
exec-brief / exec-context |
Task brief and context files, generated, never hand-built — briefs carry contract verbatim and mark embedded code as advisory |
exec-dispatch |
The dispatch act, one locked step: mints the agent ID per the registry grammar, builds the role's input file (brief+context, review package, fix package, or evidence), marks the ledger with the annotation, appends the dispatches.md row (running), and renders the role prompt bracket-clean. Nothing dispatches by hand |
exec-ladder |
The revive ladder's write path: revive/redispatch rewrites the open row's revived-rvN cell, refreshes Last-Seen, enforces EXEC_MAX_REVIVE itself, and prints the filled preamble + prompt paths |
exec-seen / exec-heartbeat |
The two liveness witnesses: refresh an open row's Last-Seen, and stamp the worker heartbeat file exec-supervise reads |
exec-prompt |
Renders any registered role template with KEY=VAL slots; refuses to emit while any [UPPER_SNAKE] bracket is unfilled — an unfilled bracket is a defect, not a style gap |
exec-review-package |
Review diffs with commit list + stat + -U10 diff in one file, per round |
exec-fix-package |
Fix-round dispatch package: verdict findings + implementer report + brief + context verbatim, so fix agents see the contract, not a paraphrase |
exec-run |
Run lifecycle in the registry: start/task/complete/pause/blocked/check |
exec-step |
The pump kernel: folds a plan run (or an initiative's phase log, or every in-flight run) and emits exactly one typed action — REPAIR-STATE/RUN-START/DISPATCH/REVIEW/FIX/REVIVE/REDISPATCH/ADJUDICATE/REPORT/GATE-STAGE/PHASE-ENTER/PHASE-GATE/CRITIQUE/ASK/WAIT/DONE. Read-only and non-improvising: "what is legal next" is a script question, not a controller judgment. Every emit carries a (wake on …) condition; repeated emits with no state change escalate to ADJUDICATE instead of looping |
exec-supervise |
Worker liveness and the revive ladder: output presence > worker heartbeat > row age, then REVIVE → REDISPATCH → ADJUDICATE, counted from the dispatch log rather than a counter file a worker could delete. Product paths are role-aware (a review row's product is its verdict, not the implementer's report) |
exec-adjudicate |
The evidence pack: assembles the adjudication set mechanically — report, latest verdict, diff, open rows, ledger and rulings tails — then dispatches SUPERVISOR. The adjudicated party does not choose the adjudicator's evidence |
exec-present |
The gate card: emits a fixed-shape summary for PHASE-GATE — entered, gate state, critique clearance, artifact paths, gate-check result. Artifact bytes never enter the controller's context |
exec-status |
The resume surface: one digest of every in-flight run and initiative with its next verb, open rows, and store drift. --write stamps .executor/RESUME.md so a cold agent reorients in one read |
exec-graph |
The graph's integrity check: every exec-step verb has exactly one actuator row, every ledger file traces to a script writer, every role template resolves, every ID grammar carries placeholders, and ledger↔plan↔dispatch↔registry references resolve. A dangling reference is a FAIL, not a NOTE |
exec-report |
Gate-on-commit: the only path a worker's result takes into run state. Refuses (writing nothing) unless the task's latest-round verdict is clean; on success appends the canonical ledger line, writes the Task status row, closes the dispatch rows, and refreshes the registry's task count — one locked act |
exec-gate |
Autonomy policy: decides whether a phase gate may clear without a human. Fail-closed — a missing, disabled, or unlisted policy reads as deny. --auto records **auto-passed** in the phase log (visibly distinct from a human passed); a pick-class phase additionally requires a scored verdict on disk, because a script can count files but cannot choose between designs |
exec-plan-regression |
Plan-set gate: resolves the initiative-level plan-regression/ dir, seeds the clearance summary, and check refuses a run while any plan is unaudited or unrepaired |
exec-critique |
The gate every authoring phase passes through: audits one component's document set against itself and its upstream contracts, then decides whether that phase may be marked passed. One script, every component — the mapping from phase to component to check catalog is a registry row, so adding a component is data, not a fork. Hardened where plan-regression was soft: verdict: PASS beside a non-zero high/medium is refused rather than believed, the three-round cap is enforced by the script, and a waiver needs both a note and an initiative ruling behind it |
exec-run check |
Drift + semantic audit: registry row vs ledger, verdict CONTENT (a FAIL verdict blocks), exact task set with latest-state reduction, final verdict lineage (a failing final re-review supersedes an earlier clean one), completed tasks present in the Task status table, branch topology (branch/task agreement, recorded merges present on the plan branch, unmerged task branches under a clean final verdict, orphan worktrees) — exit 1 names the failure |
exec-branch |
Branch lifecycle: start/merge/abandon for the plan branch (merge refused unless the review audit passes), task start|merge|abandon for task branches (merge refused unless the task's own verdict is clean; --worktree for parallel waves), side/spike for work that is neither, and status for the branch stack + orphan worktrees |
exec-evidence |
Per-criterion evidence files in the initiative's tracked verification/evidence/PNN/, immutable per round (-attempt2 on same-round reruns), atomic publish, per-round state stamp (branch, commit, dirtiness) |
exec-store-check |
Thinking-store integrity gate: registry ↔ folders ↔ Documents table ↔ frontmatter statuses ↔ cross-links ↔ evidence citations ↔ phase-log chronology, plus per-kind document contracts (architecture must carry a Mermaid diagram, specs need requirement headings + a resolving verification link, active options need a decision, criteria_count must match V<nn> rows), the doc-code boundary, the brainstorm record, branch provenance, and plan-regression clearance (execution entered without a passed/skipped plan-regression phase fails) |
exec-ruling |
Record a decision taken on the human's behalf — to the rulings log and the local decisions store. --answered marks it as a reply to a question you asked; --unsolicited "<verbatim>" records input the human offered unasked, and --stop halts the run mechanically instead of leaving the controller to decide whether to comply |
exec-scan-secrets |
Credential-shaped content scan across both stores; reports file:line, never the value |
One branch per artifact level, named after the artifact's ID
(full spec): initiative/INIT-NNNN forks from
wherever the human currently is (fork point recorded), plan/INIT-NNNN-Pnn
forks from the initiative branch, and task/INIT-NNNN-Pnn-Tnn forks from
the plan branch tip at dispatch — so task N+1 already contains task N's
merged work.
Each level merges only through the gate its level requires: a task merges
when its own review verdict is clean, a plan merges when its final verdict
exists and the audit passes, and merging the initiative onward is always
the human's explicit decision at handoff. The plan branch's history is
therefore only reviewed merges — clean to bisect, one git revert -m 1
from undoing a single task.
Parallel tasks get a worktree each (task start --worktree), because two
agents cannot share one working tree. A plan that is a single chain may
declare sequential: true to commit direct to the plan branch — the lint
requires the dependency chain to justify it. Work that is neither plan nor
task goes on side/ (merged on a recorded ruling) or spike/ (never
merged); an urgent out-of-scope fix goes on hotfix/ from the base branch.
flowchart LR
BASE["base branch"] --> INIT["initiative/INIT-0004"]
INIT --> P1["plan/INIT-0004-P01"]
INIT --> SIDE["side/INIT-0004-slug"]
P1 --> T1["task/INIT-0004-P01-T01"]
P1 --> T2["task/INIT-0004-P01-T02"]
T1 -->|"review-gated"| P1
T2 -->|"review-gated"| P1
SIDE -->|"human"| INIT
P1 -->|"final verdict + audit"| INIT
INIT -->|"human at handoff"| BASE
BASE --> HOT["hotfix/slug"]
HOT -->|"human"| BASE
Every review is a round with an ID (INIT-0004-P01-T03-R02), a diff
file, and a verdict file carrying YAML frontmatter. Findings are labelled
(C1, I2, M1), cited to the spec requirement they violate
(INIT-0004-SPEC-01-R07), and live in files a fixer reads directly — the
controller transcribes nothing. The final whole-branch review walks every
declared cross-task seam and triages every deferred or parked finding.
Re-reviews do two jobs: impact review of the fix (following what it actually affects — including unchanged callers) before finding closure, so a regression the fix introduced in untouched code is still caught, and a fix-only lens never hides it.
Briefs, contexts, ledger, rulings, preflight scan, dispatch log, reports,
verdicts, evidence files — all carry the same YAML identity block
(kind, id, initiative, plan, created_at, …), defined in the
frontmatter contract. An agent
reading any file cold knows exactly what it is holding.
The spec's verification strategy names one criterion per requirement with
its exact command. The verification phase runs each row fresh against the
current commit and reports four honest statuses: PROVEN, FAILED,
NOT-RUN, UNAVAILABLE. A single NOT-RUN blocks the word "complete" —
and nothing upgrades it by inference. Raw observed output lands in
per-criterion evidence files the outcomes table cites.
After a context loss or model switch, the controller reads the ledger — not its recollection. Completed tasks are not re-dispatched; the ledger's identity block refuses a ledger that belongs to another plan; live subagent identities are recorded so a fix round can resume rather than replace. Nothing in either store is ever deleted by a skill — pruning is a human decision.
The engine runs two axes over the same machinery, and confusing them is the main thing to get right when reading a transcript.
The phase axis is the initiative's own lifecycle: intake → discovery →
architecture → design → specification → planning → plan-regression →
execution → review → verification → handoff. Its atomic step is
exec step INIT-NNNN, which prints PHASE-ENTER or PHASE-GATE. Between
those two, exec-dispatch --role author mints and logs the phase's AUTHOR,
which writes the artifact; the phase's critique audits it.
The plan axis is what happens inside execution: exec step PLAN_FILE
prints RUN-START, DISPATCH, REVIEW, FIX, REPORT, GATE-STAGE,
REVIVE, REDISPATCH, ADJUDICATE, WAIT or DONE. Workers run on their
own branches; a worker's result reaches run state only through exec-report,
which refuses unless the latest verdict is clean.
The canonical phase order lives in exactly one place — exec_phases() in
_exec-lib.sh — and both the validator and the pump fold against it. Two
hand-maintained copies would let a run advance and refuse in the same turn.
The controller runs one loop: run exec, read the single action word it
prints, do that, go back. It never authors, never judges, never passes a gate.
The reason is not stylistic — the controller is the least reliable worker in
the system. It dies, compacts, drifts, and is under pressure at exactly the
moment a judgment call matters, so the design assumes it will get clever and
takes away the opportunity.
That includes the write side. Nothing the controller used to do by hand
survives as a hand-step: exec-dispatch performs the whole dispatch act
(mint the agent ID, build the role's input, mark the ledger, append the row,
render the prompt), exec-ladder rewrites the revive rungs and enforces the
bound, exec-seen and exec-heartbeat move the liveness witnesses,
exec-adjudicate assembles the adjudicator's evidence set, and exec-report
is the only path a result takes into run state. exec-graph check fails the
build when any of that wiring drifts — a ledger file with no script writer, a
verb with no actuator, a reference that no longer resolves.
Every step is therefore one of two things, and the distinction is load-bearing:
- A dispatch is intellectual work — authoring, auditing, implementing, reviewing, adjudicating. It goes to a named role with a registered prompt template, and its output is an artifact someone else will read.
- A script call is a transition —
exec-initiative phase … entered,exec-workspace,exec-branch merge. It writes state, and no judgment exists anywhere in the path.
The test: would a subagent reading only its prompt know more than you do right now? If yes, dispatch. If the prompt would be a transcript of the command you are about to type, run the command.
There is no ad-hoc dispatch. references/layout.md carries a registry of
roles, and every row names the prompt template that implements it — the
auditor, the repairer, the re-auditor, the implementer, the reviewers, the
supervisor, the scouts. The suite checks that every action exec-step can
emit has a decision-table row, and that every registered role's prompt exists
on disk. Adding an action or a role without wiring it fails the build rather
than a run.
The prompts themselves carry the dispatch contract: an identity block, a self-critique pass, a verification section, and a machine-readable return block. They also carry the states a naive agent invents behaviour for — a missing upstream artifact, two inputs that contradict, an output file left behind by an aborted prior run, a test that fails for a reason outside the task's remit. A prompt is the only definition of what a subagent does, so an undefined state is an unreviewed behaviour, not a style gap.
The critique engine is the extension point worth knowing about. Which
artifacts a component owns, where its clearance record lives, and which check
catalog applies are one row of data in exec_critique_components():
component | phase | set_spec | summary_dir | catalog
Adding a twelfth component means adding a row and a catalog section in
executor-critique/SKILL.md. It does not mean writing a script, forking
exec-plan-regression, or teaching the pump a new case. plans is the proof:
it resolves to the pre-existing plan-regression/ home and reads the same
summary.md the shipped skill writes — one gate with two front doors, not two
gates that can disagree.
The registry is positional data, so its shape is enforced where it is read:
exec-critique counts each row's fields and refuses a miscount. A row with one
field too many reads as plausible-but-wrong, and that class of bug once put two
components' clearance records in a directory literally named -.
flowchart LR
subgraph THINK["docs/executor/ - tracked"]
C["Charter"] --> R["Research, Options"]
R --> A["Architecture, ADRs, Interfaces"]
A --> D["Design"]
D --> S["Spec"]
S --> P["Plans"]
end
subgraph EXEC[".executor/ - git-ignored by default"]
L["Ledger, Rulings"]
B["Briefs, Contexts"]
RP["Reports"]
V["Diffs, Verdicts"]
EV["Evidence files"]
PR["Plan-regression audits, fixes, gate summary"]
end
P -->|"plan-set audit"| PR
PR -->|"clearance"| B
P -->|"each task dispatch"| B
B --> RP
RP --> V
V --> L
docs/executor/— the thinking record. Git-tracked: charter, research, architecture, decisions, interfaces, design, spec, risks, plans, and the verification strategy + outcomes ledger. A reader who clones the repo gets the complete reasoning..executor/— the execution record. Git-ignored by default, safe to commit if you choose: task briefs and contexts, implementer reports, review diffs and verdicts, evidence files, the progress ledger, rulings. Never deleted — the reasoning is the point.
The split is durability-of-audience, not durability-of-value. .executor/
resolves from the main repository root, so removing a worktree cannot
destroy the execution record.
INIT-0004 the initiative
INIT-0004-CHTR-01 its charter
INIT-0004-RSCH-02 a research note
INIT-0004-SPEC-01-R07 requirement 7 inside that spec
INIT-0004-P01 a plan
INIT-0004-P01-T03 task 3 of that plan
INIT-0001-P01-T03-R02 review round 2 of that task
Addressable requirements are what let a review finding name the exact contract it violates, and what lets a plan task declare precisely which requirements it discharges.
.executor/ may be committed, so everything in both stores is written as
though it will be public. Credentials, tokens, and personal data never go
into any Executor artifact — a redacted existence statement and a safe path
instead. skills/executor/references/safety.md defines the required scan
before any handoff, and reviewers treat credential-shaped content in any
diff as a Critical, stop-and-tell-the-human finding.
| Doc | Purpose |
|---|---|
| CONTRIBUTING.md | How to change skills, references, and scripts |
| CODE_OF_CONDUCT.md | Participation standards and enforcement |
| SECURITY.md | Reporting vulnerabilities and how artifact secret-hygiene works |
| CHANGELOG.md | Notable changes, newest first |