Skip to content

Latest commit

 

History

History
314 lines (284 loc) · 17.4 KB

File metadata and controls

314 lines (284 loc) · 17.4 KB

herdr-guard — Spec (v1, runtime 0.2.0)

Cross-agent command policy layer for Herdr. Watches herdr panes for dangerous commands, then audits, alerts, or interrupts — from one policy, regardless of which agent or shell is running in the pane.

Honest capability model (verified against herdr 0.7.5 + red-teamed)

Coverage depends on pane type. This table is the product's contract — it belongs in the README verbatim:

Pane type What the guard sees interrupt behavior
Interactive shell (zsh/bash, canonical echo) Everything typed, including unsubmitted input (tty echo) Best-effort pre-execution Ctrl+C request; request acceptance is recorded and prevention remains unknown
Shell in raw/no-echo mode (stty -echo, curses wrappers) Nothing typed No request. stty -echo itself is an alert-class rule
TUI agent (Pi, Claude Code, Codex) Only what the TUI renders — permission prompts, expanded tool views. Tool-executed command strings are usually NOT echoed Usually no request. Do not rely on it
Plain process output (logs, builds) Everything printed No request unless the pane is classified as a shell
Popup panes Nothing — popups have no pane id, no pane events, no pane API No request. Hard blind spot, documented

The guard matches command text, not command intent. Semantic obfuscation (base64 -d | sh, r''m, $x -rf, python shutil.rmtree) defeats content matching; obfuscation-indicator alert rules make attempts loud but can't stop them. Harness-level hooks remain the enforcement point inside TUI agents — and the guard now ships that path: harness reporters (see "Harness reporter ingest" below) submit tool calls pre-execution for a unified audit + policy verdict the harness can enforce.

Architecture

Stack: plain ESM Node.js (>= 20), zero runtime deps, bash shims in the manifest. Same stack as herdr-swarm / herdr-browser.

herdr-guard/
  herdr-plugin.toml
  scripts/*.sh            # manifest entrypoint shims
  src/
    watcher.mjs           # long-running guard process (the Guard pane)
    herdr-socket.mjs      # NDJSON socket client: connect, request, subscribe, reconnect
    policy.mjs            # pure policy engine: load/merge config, validate, match -> decisions
    rules-default.json    # shipped default policy (seeded into config dir)
    audit.mjs             # JSONL append, redaction, sanitization, partitioned rotation
    render.mjs            # dashboard rendering (ANSI, sanitized)
    reporter.mjs          # harness reporter ingest: NDJSON unix-socket server
  hooks/
    claude-code-pretooluse.mjs  # shipped Claude Code PreToolUse reporter (fail-open)
  tests/*.test.mjs        # node:test — policy engine + socket client (fake NDJSON server)

Harness reporter ingest

The one place prevention is actually possible: an agent harness reports each tool call BEFORE execution and can honor the verdict.

  • Transport: NDJSON request/response over a unix socket at the well-known per-user path $XDG_STATE_HOME/herdr-guard/reporter.sock (default ~/.local/state/herdr-guard/reporter.sock) — deliberately NOT the per-session herdr state dir, because reporters run inside agent processes without herdr's plugin environment. Dir 0700 (created only if missing — a user-overridden path never gets its existing parent chmodded), socket 0600. Claiming is race-safe: an atomic (O_EXCL) pid lock file gates the unlink-and-bind, a dead holder's lock is reclaimed, and close() removes only a socket/lock the instance owns — a losing guard's shutdown can never delete the surviving guard's live socket. A live socket (another guard) is left alone and logged. Override with HERDR_GUARD_REPORTER_SOCKET (empty means unset, on both the guard and hook sides).
  • Request: {v:1, kind:"tool_call", agent, tool, command, cwd?, session?}. Response: {ok, verdict: "deny"|"warn"|"allow", enforcement, rule_id, severity, reason}. Mapping: interrupt→deny, alert→warn, audit/none→allow.
  • Matching runs with paneType: "harness" so prompt_only never gates a raw reported command (there are no prompt glyphs to find). Project overrides merge by the reported cwd. Multi-line commands take the worst-line verdict.
  • Every match is audited (source: "harness:<agent>", decision advise-deny / advise-warn / log-only, prevention: "unknown" — the guard cannot observe whether the harness honored the verdict). Dedupe/rate-limit/notification-coalescing reuse the pane pipeline with a synthetic harness:<agent> pane key; interrupt-tier is never suppressed. pause yields allow + enforcement: "paused" while still auditing — same contract as panes: pause stops actions, never the record.
  • Shipped reporter: hooks/claude-code-pretooluse.mjs, a zero-dependency Claude Code PreToolUse hook. deny → permissionDecision: "deny", warn → "ask", allow → silent. Strictly fail-open (500ms deadline, exit 0 on any failure): a broken or absent guard must never break the harness. Residual: killing the guard silences this path; pane-side tamper rules make that loud.

The Guard pane (watcher)

One long-running [[panes]] entrypoint (placement = "split"). Lifecycle:

  1. Bootstrap: session.snapshot → live panes (excluding own plugin panes by plugin id, never by process name).
  2. Subscribe-first, reconcile-after (per pane): open one base pane.output_matched subscription with a combined alternation regex and at most one replaceable project-policy subscription (source: "recent_unwrapped", lines: 5, strip_ansi: true), THEN pane.read for a baseline. Cwd and policy transitions acknowledge a replacement project stream before atomically swapping ownership and closing the previous stream; a rejected replacement is audited, retried, and leaves the old stream live. No-pattern transitions intentionally close the prior project stream. Replay suppression is content-based: drop events whose matched lines are a subset of the post-subscribe read; the 500ms-arrival window is only a fallback heuristic. Never act on suppressed events — an interrupt replay would ctrl+c whatever the pane is doing now.
  3. Global subscriptions: pane.created (async subscribe-first per step 2 — never block the event loop on baselining), pane.closed / pane.exited (cleanup + post-mortem sweep: final pane.read of the dying pane, locally matched, so split-run-close attacks at least get audited after the fact). Events arriving with the ack are baseline, not transitions.
  4. On match event: herdr-side matching is edge-triggered, so the guard re-scans the event's read.text locally against all rules. Dedupe is: time-bounded (suppress identical repeats only within ~2s), LRU-capped (256 entries/pane), keyed on pane, line, and rule id — and interrupt-class rules are never deduped across interrupt events (a rule severe enough to ctrl+c once is severe enough to ctrl+c twice; this kills the pre-seeding attack).
  5. Sweep backstop: pane.read per pane, scheduled per-pane since-last-sweep (~10s each, bounded concurrency — NOT a global round-robin cycle that degrades with pane count), locally matched + deduped. Push path is the primary record; eviction racing the sweep is a documented limitation.
  6. Rate limits: max 10 actions/min/pane; beyond that, one coalesced "N suppressed matches for rule X" audit entry per minute. Notifications coalesce by rule id. Flooding cannot evict real history: severity- partitioned files — audit.interrupt.jsonl rotates separately from audit.jsonl (10MB × 3 generations each, 0600 preserved).
  7. Decide and act (highest matching severity wins):
    • all match tiers record decision, interrupt_request, and prevention: "unknown" alongside pane/rule/process context;
    • interrupt_request is exactly accepted, failed, or not-requested; accepted means only that the Herdr RPC succeeded;
    • alert / interrupt: request notification.show (best-effort; the audit entry is the reliable local record);
    • shell interrupt: request pane.send_keys Ctrl+C to the event's own pane_id only. Pinned invariant: pane_id comes from the event, never from matched text or config. TOCTOU acknowledged: the request targets whatever is in the pane then, and Guard cannot observe prevention.
  8. prompt_only matching: rules with prompt_only: true (DEFAULT ON for interrupt rules) only fire when the matched line looks like a prompt line (leading ❯/$/%/╰─ glyph). Prevents ctrl+c-ing vim, less, build logs, and cat RUNBOOK.md — and prevents FP-storm alert-fatigue attacks that train users to pause the guard. Trade-off: heredoc bodies and script text don't trigger interrupt; they still audit.
  9. Enrichment: pane.process_info for foreground argv/cwd; classify pane_type from process name + terminal title. Override cache is keyed by pane cwd and invalidated on cwd change, not just mtime. Project policy changes replace their owned subscription instead of accumulating streams.
  10. Fail-visible: on initial server absence or a later socket drop, the dashboard reports disconnected state; reconnect uses backoff and then a full bootstrap (snapshot → subscribe-first → reconcile). Bootstrap work is generation-gated: newer reconnect/config/retry attempts supersede older attempts, and only the newest generation may own streams or publish ready.
  11. Self-healing watchdog: [[events]] hooks on pane.closed AND pane.exited — if the Guard pane died, the hook auto-reopens it via plugin.pane.open and notifies. A sibling-session check enumerates ~/.config/herdr/sessions/*/herdr.sock and alerts on sessions the guard isn't watching. Residual (documented, unfixable in v1): herdr plugin disable, pkill, herdr server stop — the guard is advisory against agents with socket access. Upstream ask: socket ACL / read-only token, popup visibility in pane API.
  12. Dashboard: policy state (active/paused/disconnected), panes watched, rules loaded plus actual load-warning count, matches this watcher run, last N audit entries. ANSI clear+reprint on event/resize. All event-derived strings stripped of C0/C1/OSC/U+2028/U+2029 before rendering — and the same sanitization applies to audit-log writes.

Policy config

  • Main: $HERDR_PLUGIN_CONFIG_DIR/rules.json. Seeded from shipped rules-default.json on first run or via guard reset-rules (backup first). State dir 0700; audit + rules files 0600.

  • Format:

    {
      "version": 1,
      "enforcement": "active",
      "allow_project_override": false,
      "rules": [
        {
          "id": "rm-rf-root",
          "severity": "interrupt",
          "match": "regex",
          "prompt_only": true,
          "pattern": "rm\\s+(-[a-zA-Z]*r[a-zA-Z]*f|-[a-zA-Z]*f[a-zA-Z]*r)[a-zA-Z]*\\s+(/|~|/\\*|\\$HOME)",
          "reason": "Recursive force delete near filesystem root"
        }
      ]
    }

    match is regex (RE2-safe — herdr evaluates with the Rust regex crate) or substring.

  • Rule rules: no ^ anchors (matched lines carry prompt glyphs); validator at load: cap length 512, reject backrefs, reject unparseable regex (log + notify on rejection).

  • Default rules (from pi damage-control, pi-library sp-damage-control

    • safe-mode, red-team additions; every rule has hit/near-miss coverage in tests/rules-default.test.mjs):
    • interrupt: rm -rf rootish paths, dd of=/dev, mkfs, wipefs/blkdiscard/shred on devices, shell redirects onto block devices, recursive chmod/chown on rootish paths, find / -delete, fork bombs, crontab -r, terraform destroy, kubectl delete prod-ish contexts
    • alert: sudo, curl|sh / wget|sh, cat .env* / SSH-key and credential-store reads / security find-generic-password -w, npm publish and the wider publish family (cargo publish, twine upload, gem push, yarn/pnpm publish), aws s3 rm|rb|sync --delete, AWS/GCP/Azure resource deletion, PaaS app destruction, DB DROP/TRUNCATE (prompt-only), kubectl delete namespace / helm uninstall, docker system prune -a and volume removal, git push --force|--mirror|--delete (force-with-lease is audit-only), gh repo delete, firewall disabling, exfil (scp|rsync of ~/.ssh, ~/.aws, ~/fsw-bid-data; curl uploads of secret material), guard tampering (herdr plugin disable, killing Herdr, deleting guard rules/audit files), evasion indicators: stty -echo, stty raw, tmux.*(-d|-b), screen -dm, disown, setsid, | at now, history clearing (history -c, HISTFILE=/dev/null, unset HISTFILE), base64 -d / xxd -r / printf '\x..' piped to shell, eval $(, sh -c "$(
    • audit: git destructive variants (clean -f, checkout -- ., branch -D, filter-branch/filter-repo, stash drop|clear, push --force-with-lease), generic rm -rf
  • Project override: <workspace cwd>/.herdr-guard.json, merged lazily per pane cwd.

    • May add rules (substring only — repo-controlled regex never reaches the herdr server) and may raise severity to alert at most (raise-to-interrupt is self-DoS with no defensive value).
    • May never disable or lower config-dir rules unless "allow_project_override": true. Every applied override is audit-logged with its workspace path.
  • Writes are atomic (tmp + rename). On parse error: keep last-good config + alert. Never fall back to empty.

Enforcement changes are privileged events

pause / resume flip enforcement (watcher mtime-polls, 2s). Every enforcement flip, config reload, and rule-set change is audit-logged (before/after counts) AND fires notification.show. Watcher reloads emit config-change-detected before transport reset, then config-change-applied after the winning bootstrap is ready or config-change-failed while retaining a pending retry. reset-rules records bounded old/new rule counts and requests a notification after reseeding. pause takes an optional TTL (default 15min) after which enforcement auto-resumes.

Audit log

$HERDR_PLUGIN_STATE_DIR/audit.jsonl + audit.interrupt.jsonl, 0600, 10MB × 3 generations, partitioned so floods can't evict interrupt history. Before writing, every string field is truncated to 200 chars, redacted (KEY=VALUE, bearer/sk-/ghp_ token shapes), and sanitized (C0/C1/OSC/U+2028/U+2029). Match records separate the policy decision from the interrupt request result and always report prevention: "unknown". The audit log is sensitive data; README says so.

Manifest surface

  • [[panes]] guard (split) — watcher + dashboard.
  • [[actions]] open — open/focus the Guard pane.
  • [[actions]] pause [ttl] / resume — enforcement control (audited).
  • [[actions]] test — popup prompt: type a command string; shows matching rules + verdict, using realistic prompt-decorated lines (dry-run).
  • [[actions]] reset-rules — reseed config (backup first).
  • [[startup]] — seed config if missing, then idempotently plugin.pane.open the guard entrypoint (check session.snapshot first). Without this the guard watches nothing until manually opened.
  • [[events]] on = "pane.closed" and on = "pane.exited" — self-healing watchdog (reopen guard pane + notify).

Known limitations (README material, accepted for v1)

  • herdr plugin disable / pkill / herdr server stop — guard is advisory against socket-privileged agents. Upstream ask: socket ACL.
  • Popup panes are invisible execution channels. Upstream ask: popup API.
  • Semantic obfuscation defeats content matching.
  • Sibling named sessions run separate sockets (guard alerts on them, doesn't watch them).
  • Raw-mode/no-echo shells hide typed input (alert rule on stty -echo is the mitigation).
  • Nested-mux detached execution (tmux new-window -d, screen -dm, nohup+disown) — launcher line may alert; execution is invisible.
  • Interrupt TOCTOU: a Ctrl+C request targets whatever is currently foreground; accepted delivery and prevention are not observable.
  • Scrollback eviction can beat the sweep; push path is the primary record.
  • Smoke-test matrix required before publish: user scrolled up in scrollback, terminal resize, OSC8/DCS/APC/bracketed-paste ANSI correctness of strip_ansi, named sessions.

Testing

  • node:test for policy.mjs: load, merge precedence, override add/raise-capped-at-alert enforcement, pattern validator, severity routing, redaction/sanitization, prompt_only gating, rate limiting, time-bounded dedupe incl. interrupt no-dedupe.
  • herdr-socket.mjs against a fake NDJSON server: bootstrap, subscribe-first/reconcile, content-based replay suppression, edge-trigger local re-scan, post-mortem sweep, reconnect → re-bootstrap.
  • Safe manual smoke: confirm Herdr 0.7.5, perform two sequential read-only session.snapshot RPCs, run the policy demo without executing its input, and inspect that an interrupt dry-run says it would request Ctrl+C while prevention remains unknown. Live key delivery is optional and must not use a destructive command.

Publish

Repo StructuPath/herdr-guard, topics herdr-plugin, herdr, security. MIT license. README leads with the coverage-matrix table and the known- limitations list. Honesty is the differentiator.