A safe, read-only CLI for querying /agent observability surfaces — built for AI agents
(Claude Code and friends) that need to ask a running service "what are you doing right now?"
without SSH, database clients, or the ability to break anything.
A convention, not a framework: a service mounts a token-gated, read-only observability API
at /agent. GET /agent is a self-describing index of the endpoints it offers:
{
"endpoints": [
{ "path": "/agent/version", "description": "build SHA and start time" },
{ "path": "/agent/health", "description": "liveness and dependency checks" },
{ "path": "/agent/queues", "description": "in-memory queue depths" }
]
}Response shapes are not stable contracts — the consumers are your own agents, updated alongside the services. agentctl is therefore deliberately dumb transport: fetch, pretty-print JSON, exit codes. No per-service schema parsing.
agentctl is what curl looks like after you delete everything an agent could misuse:
- GET-only by construction. There is no method flag. The only HTTP verb in the codebase is
a single
http.MethodGetliteral (enforced by a test that greps the sources). - Path confinement. Requests can only go to paths under
/agent/on base URLs registered in the config file. No--urlflag exists. Traversal (.., encoded variants) is rejected. - Token hygiene. Tokens live in one config file, never on the command line. Internally they
are wrapped in a
Secrettype whose every formatting path (fmt verbs, JSON) yields a fingerprint (tok:1a2b3c4d), and all diagnostics are scrubbed before printing. - Redirect pinning. Max 3 hops, same scheme+host+port as the registered base URL, and the
target path must stay under
/agent. Anything else is a transport error. - Sane timeouts (10s default, 8s for
status), 10 MiB response cap, no prompts, no color.
go install github.com/TysonLabs/agentctl@latest- Write
~/.config/agentctl/services.toml(chmod 600):
[payments.dev]
base_url = "https://dev.example.com"
token = "at_xxxxxxxxxxxx"
[payments.prod]
base_url = "https://pay.example.com"
token = "REPLACE_ME" # placeholder → listed as "not wired", never called
[payments.meta] # informational only — shown by ls, never fetched
repo = "org/payments"
unit = "payments.service"- Explore:
agentctl ls # what's registered, what's wired
agentctl endpoints payments.dev # what the service offers
agentctl get payments.dev version # any of: version, /version, /agent/version
agentctl get payments.dev "logs?limit=20"
agentctl status # /agent/version + /agent/health across all wired servicesagentctl exists so that coding agents can operate a fleet of services with curl on their deny list. The workflows below are the ones it was built around; the service names and endpoints are illustrative.
An agent is handed "checkouts are slow since noon" and has to come back with evidence, not a restart. The loop:
agentctl endpoints shop.prodto see what the service can answer. Trust the live index over any runbook; surfaces evolve.- Baseline with
version(did it just restart? does the SHA match the last deploy?) andhealth(pools, scheduler last-runs, per-dependency last success/error). - Sample in-memory state twice, a minute apart, to tell a growing queue from a stable one.
- Pull logs filtered server-side: a
sincebounded to the symptom window andq=terms taken from the symptom. Start narrow, widen only if empty. - Report a diagnosis and a recommended action for a human to take. Triage never mutates.
If /agent itself is unreachable, exit code 3 is the finding: the process is down or the network
path is broken, and the agent says so instead of guessing.
After a push, nothing counts as verified until agentctl get shop.prod version reports the
deployed SHA. Then health is checked for anything the deploy degraded, and only then does the
agent exercise the changed behaviour. A health payload can also carry "restart owed" style
fields, so a config change the process could not hot-apply shows up here rather than a week later.
agentctl status gives one line per wired service and environment, hitting version and
health with an 8s timeout. Exit 2 means at least one service answered with an HTTP error, exit
3 means at least one was unreachable. A 502 from a reverse proxy in front of a dead process
shows up as FAIL HTTP 502 on /agent/version rather than as a hung command.
Two services that talk to each other in production do not need to talk to each other for observability. When one records a problem report, it stores only a masked correlation id and prints the follow-up for the triager:
agentctl get upstream.prod "logs?q=<correlation_id>"
The human or agent doing triage performs the join by hand through agentctl. The services stay decoupled, sensitive data stays out of the reporting service's database, and the observability path never becomes a runtime dependency.
Surfaces built for this workflow have converged on roughly the same set, whatever the language:
| Endpoint | Answers |
|---|---|
/agent/version |
What build is running, since when, on which host? |
/agent/health |
Are dependencies healthy? Pools, schedulers, last success/error per integration. |
/agent/state |
What is in memory right now? Queue depths, connection counts, oldest-entry age. |
/agent/config |
What is the resolved effective configuration, with every secret redacted? |
/agent/logs |
What was logged recently? An in-memory ring, filterable server-side. |
/agent/<thing>/<id> |
Detail on one object, addressable only by an id the caller already holds. |
The last row matters: detail endpoints keyed by an unguessable reference (an incident number a user was shown, a capability id a session owner minted) let a surface expose depth without letting anyone enumerate users or sessions.
- Location:
--config PATH>$AGENTCTL_CONFIG>~/.config/agentctl/services.toml. - Each
[service.env]table needsbase_url(http/https, no userinfo/query/fragment) andtoken. - A
[service.meta]table is informational (repo, unit, owner, …) — shown byls, never fetched. - Placeholder tokens (
REPLACE_ME,CHANGEME,TODO,…,<...>, all-x, anything under 8 chars) mark a service not wired:lsshows it with the reason,get/endpointsrefuse it,statusskips it. - Keep the file
chmod 600; agentctl warns (but proceeds) if group/other bits are set.
| Command | Behavior |
|---|---|
agentctl ls |
list services/envs, wiring status, base URLs (never token material) |
agentctl get <svc.env> <path> [--raw] |
GET under /agent/; pretty-print JSON, --raw for bytes |
agentctl endpoints <svc.env> |
fetch GET /agent and render the descriptor table |
agentctl logs <svc.env> [log flags] |
GET /agent/logs as one readable line per entry (see below) |
agentctl status [svc.env ...] |
fan out /agent/version + /agent/health, one line per service |
agentctl version |
print agentctl's own version |
Global flags: --config PATH, --timeout DUR (default 10s; status default 8s), --help.
| Code | Meaning |
|---|---|
| 0 | success (status: all wired queried services 2xx on both endpoints) |
| 1 | usage error, config error, unknown service.env, not-wired target, rejected path |
| 2 | HTTP status ≥ 400 (body still printed to stdout) |
| 3 | transport: DNS/dial/TLS/timeout, refused redirect, body over cap |
| 4 | logs --wait: no matching entry before the deadline |
Output is designed for LLM agents: stdout is the answer only; stderr carries one-line
agentctl:-prefixed diagnostics. No color, no spinners, no prompts.
agentctl logs recursivecx.prod --level warn --since 30m # recent warnings, readable
agentctl logs vector-dialer.prod --q "dial-customer" --limit 50
agentctl logs recursivecx.prod --q "E911 link" --wait 20m # exit 0 when it shows up, 4 if it doesn'tIt reads the entries list from /agent/logs and prints one line per entry:
time LEVEL [source] message | key=value .... The three shapes services emit today
(ts/target/message, time/message/fields, dt/msg/fields) render the same
way, and --json prints the normalized entries instead. Newlines are flattened and terminal
controls (C0, DEL, C1, and Unicode bidirectional controls) are dropped, so log content can't
inject escapes or visually reorder an entry. Configured bearer-token values are scrubbed.
--q (substring), --level (minimum severity) and --limit are passed to the service,
which does the filtering. --since takes a duration back from now (30m) or an RFC 3339
time. --wait DUR polls every --interval (default 15s) until an entry matches, counting only
entries newer than the start of the wait (or --since). Without an explicit --since, an
initial read uses the service's HTTP Date to account for clock skew; subsequent polls keep
that resolved since fixed. It retries transport errors and 5xx, which happen mid-deploy, and
fails fast on any other HTTP error. The wait duration is an overall deadline, including time
spent in requests. It's still only GET requests under /agent/.
Capabilities that do not exist: non-GET methods, arbitrary URLs, custom headers,
--insecure, request bodies, tokens on the CLI, config-write commands (a wire/add
command would put tokens in shell history).
What does exist: bearer auth from a 600-mode file, fingerprint-only token rendering, output scrubbing, host+path-pinned redirects, timeouts, and a response size cap.
## Observability via agentctl
- `agentctl ls` — services you can query; only "wired" ones are callable.
- `agentctl endpoints <svc.env>` — discover what a service exposes.
- `agentctl get <svc.env> <path>` — read-only GET under /agent; pretty JSON on stdout.
- `agentctl status` — quick fleet health; exit 0 = all good, 2 = HTTP errors, 3 = unreachable.
- It cannot mutate anything: GET-only, /agent-only, registered hosts only.Non-normative conventions that make a surface pleasant to consume:
- Bearer-token auth; the surface only registers when a token is configured.
- Read-only forever; mutations belong elsewhere with their own auth.
- Cheap by construction: in-memory state, O(1) lookups — nothing a caller could use to load you.
- No secrets, no customer PII in responses or logs.
GET /agentreturns{"endpoints":[{"path":...,"description":...}]}so tools and agents can discover everything else. Write each description as the question it answers ("What build is running?") — that is what an agent reads when deciding where to look.- Off by default: an empty token means the route group is never mounted, not "mounted but 401".
- Make index drift impossible: either drive the router from the same table that renders the index, or add a test that fails when a registered route lacks a descriptor.
- Sanitize at the point of capture (log ring, event cache) rather than at serve time, so a query parameter can never become a search oracle for the raw value. Let opaque ids survive masking; they are the join keys triage depends on.
- When a field can legitimately be unknown, return why (
"not_configured","bypassed",{"configured": false}) instead ofnull. A bare null costs someone an hour later. - Bound every response: default and maximum
limit, per-entry byte caps, and a total that stays well under agentctl's 10 MiB ceiling.
agentflow lives in the same repo as a separate binary (cmd/agentflow), so agentctl
keeps its read-only guarantee and permission allowlists stay per tool. An import-boundary
test keeps the two apart. Each agentflow command does one job and reports a JSON result
and an exit code. It is a set of tools, not a harness: the workflow itself stays in prose,
in AGENTS.md.
agentflow codex --base main # codex's built-in reviewer over main...HEAD
agentflow codex --base main --prompt-file brief.md # your brief, with the scoped diff inlined
agentflow codex --uncommitted --prompt-file brief.md --path internal/flow # one area at a time
agentflow codex --prompt-file plan-review.md --dir ~/src/repo # any read-only task
agentflow codex --base main --prompt-file fix.md --write # fix mode (workspace-write)What it guarantees, each one a way a hand-typed Codex invocation has failed:
| Failure | What agentflow does |
|---|---|
Hangs forever on Reading additional input from stdin... |
stdin is never inherited: it is /dev/null, or the prompt file (read to EOF) |
The review subcommand has no --sandbox flag and inherits workspace-write |
sandbox always pinned via -c sandbox_mode=...; read-only unless --write |
| Scope flags can't be combined with a custom prompt (clap error) | a scope plus a prompt inlines the scoped diff (git diff -M, untracked files included) |
| Runs for an hour, or wedges silently | --timeout (default 40m) and --stall (default 10m with no --json events and no growth of codex's session log) kill the whole process group |
| Exit 0 with no answer read as "clean review" | ok requires exit 0, no failed turn, and a non-empty final answer |
| Giant prompts stall in reasoning | prompts over --max-prompt-bytes (default 80000) are refused with a hint to split by --path |
-C outside a git repo dies on the trust check |
--skip-git-repo-check is added only when --dir is not a git work tree |
| An empty scope "passes" | an empty diff is an error, never a review |
Output: JSON on stdout (also <out>/result.json), with status, codex_exit, duration_s,
thread_id, usage, error, and paths to final.md (the answer), prompt.md,
events.jsonl, stderr.log and codex's session rollout. Runs take minutes, so agents
should start it in the background and read final when it exits.
| Exit | Status | Meaning |
|---|---|---|
| 0 | ok |
final answer written |
| 1 | — | usage or precondition error (bad flags, empty diff, prompt too large, agent missing) |
| 3 | codex_failed / claude_failed |
the agent exited non-zero or reported a failed turn |
| 4 | no_answer |
the agent exited 0 without a final answer |
| 5 | rate_limited |
usage or rate limit: wait, then retry |
| 124 | timeout |
killed at --timeout |
| 125 | stalled |
killed after --stall with no activity |
| 130 | interrupted |
agentflow was interrupted; the agent was killed |
Set AGENTFLOW_CODEX to use a codex binary other than the one on PATH.
Every prompt-driven agent is told it is a sub-agent: do the one task, report, and stop,
and don't start other agents or reviews, commit, push or open PRs. Codex's native
exec review mode has no prompt file; it uses its built-in reviewer instructions.
Without the note, a prompted reviewer that reads the repo's AGENTS.md may try to
start a review of its own fixes.
agentflow claude --base main # default review brief, diff inlined
agentflow claude --base main --prompt-file brief.md --write # fix mode
agentflow claude --uncommitted --prompt-file brief.md --max-budget-usd 3It takes the shared flags, guarantees and exit codes of agentflow codex, so a change
Codex wrote can be reviewed by Claude. The differences:
- Sandbox by tool list. Claude Code has no flag like codex's
sandbox_mode, and in headless mode it inherits your settings' permission mode, which can bebypassPermissions. agentflow runs it with--restricted(no shell or code-running tools; user, project and local settings and hooks ignored; file tools confined to--dir) and--strict-mcp-config(no MCP servers). Read-only gets exactlyRead,GrepandGlob. That's an allowlist, because the default set also has tools that create worktrees, message other sessions or upload content. Fix mode (--write) addsEdit,WriteandBash, with Bash in Claude Code's sandbox: writes only under--dir, no network, and no unsandboxed escape. (A fullgo testmay still fail there, because Go's build cache is outside--dir; Codex's workspace-write has the same limit.) - No built-in reviewer. A scope with no prompt gets a short default review brief, with the diff inlined.
- The answer comes from the stream.
okrequires exit 0 and a finalresultevent withsubtype: "success"andis_error: false; its text becomesfinal.md. An error result (error_max_turns, a budget stop, an API error) isclaude_failed(exit 3), orrate_limited(exit 5). - Cost. The JSON adds
cost_usd, and--max-budget-usdstops a run at that spend. The exit code is reported asclaude_exit, androlloutis the session transcript under~/.claude/projects/(or$CLAUDE_CONFIG_DIR).
Set AGENTFLOW_CLAUDE to use a claude binary other than the one on PATH.
agentflow ship verify recursivecx.prod --sha "$MERGE_SHA" # safest: wait for that exact build
agentflow ship verify recursivecx.prod --sha "$MERGE_SHA" --contains --repo . # accept a later forward deploy
agentflow ship verify vector-dialer.dev --sha d63a514 --once # one check, no waitingIt reads /agent/version through agentctl ($AGENTCTL, then PATH, then ~/bin), so
agentflow never holds service tokens and makes no HTTP calls itself. It replaces the
hand-written until …; sleep loops agents write after every merge.
- The commit comes from
git_commit,commit, orgit_sha, or from aGit Commit: <sha>line in theversionbuild banner. Genericsha,revision, and unrelated string fields are not trusted. Conflicting recognized values fail closed. - A short SHA matches a full one (prefix either way, at least 7 hex digits), because builds usually stamp short SHAs.
--shatakes a hex SHA (compared as given, so it works from any directory, even for a commit not fetched yet) or a git rev (origin/main,HEAD, a tag) resolved in the local--repocheckout (default: the current directory). Nothing is fetched, so after a merge prefer the merge's SHA over a remote-tracking ref that may be stale.- Exact matching is the default. With an explicit
--repo,--containsalso accepts a newer deploy that contains the expected commit ("match": "contains"). It is intentionally opt-in because a pre-rollback build also descends from the older commit being restored. - Transport failures and rollout-like HTTP errors (404, 408, 425, 429, and 5xx) are
retried until
--timeout(default 40m, checking every--interval, default 30s). Other HTTP errors and a version with no recognizable commit fail at once.
| Exit | Status | Meaning |
|---|---|---|
| 0 | deployed |
running the expected commit, or (with --contains) a descendant |
| 1 | agentctl_error / — |
usage error, or a non-retryable agentctl/HTTP failure |
| 2 | not_deployed |
--once only: not yet |
| 3 | unreadable |
/agent/version has no recognizable commit |
| 124 | timeout |
never matched before --timeout; running/error show the last state |
| 130 | interrupted |
interrupted |
agentflow worktree done feat/my-change # by branch, path, or worktree dir name
agentflow worktree done feat/my-change --dry-run
agentflow worktree sweep # list every removable worktree
agentflow worktree sweep --yes # remove them, prune stale recordsAgents tend to clean up with git worktree remove --force and git branch -D. That
works, and it also deletes uncommitted work and unmerged branches without a word.
done proves removal is safe first, and refuses with every reason if it isn't:
- Merged: the tip is contained in the freshly fetched target (
--into, defaultorigin/HEAD), or GitHub shows a PR merged into that target whose branch and head are exact matches, so squash and rebase merges count. - Clean: no modified or untracked files. Ignored build output (
target/,node_modules/) goes with the worktree, so--forceis never needed. - Unused: not locked (the lock reason and whether its owner pid is alive are
shown), no process has its working directory inside (via
lsof), and no registered worktree or other Git repository is nested beneath it. Initialized submodules are also refused because Git cannot remove their worktree without--force. - Not the repository's main working tree.
Then it removes the worktree, deletes the local branch only if it still points at the
proven head, and deletes the remote branch only when a merged PR from that same
repository had exactly that branch and head. Long-lived branches (the target, main,
master, develop, development, staging, production, release/*, hotfix/*)
are never deleted. sweep runs the same checks on every worktree; with --yes it
removes those that pass and prunes the records of worktrees whose directories are gone.
--keep-remote leaves remote branches alone.
JSON on stdout (per worktree: ok, merged_via, refusals, keep_branch, and after
removal freed_bytes, branch_deleted, remote_deleted). Sweep continues after an
individual worktree error, records it in that entry's error, prints the complete
JSON result, and exits 3. Exit codes: 0 removed (or would be, or sweep finished) ·
1 usage · 2 refused · 3 git/gh error. A missing or incomplete lsof check is a
safety refusal.
Color/TTY niceties, retries, response caching, keychain integration, --json listing output,
shell completions, config-write commands, per-service schema rendering. PRs adding request
capabilities beyond GET-under-/agent will be declined on principle.
make all runs vet, race-enabled tests, and the build. The test suite includes a source guard
that fails if any mutating HTTP verb appears in non-test code — keep it that way.
MIT © Tyson George