Interactive CLI agent runtime backed by Vertex AI Gemini 3.x — deliberately minimal, sandboxed, and usable both interactively and headlessly.
Released. Install with
brew install nlink-jp/tap/gem-agentor from the releases page (Developer ID signed, Apple-notarized, macOS arm64) — the releases page is the authority on the current version, so this line does not repeat it. It is a cli-series tool: interface stability is a promise, and breaking changes go through the org's breaking-change process (ADR-0061). See the RFP for the full specification.
gem-agent is an independent agent runtime: a minimal, auditable agent loop
(read / edit / shell / MCP / approval) on Vertex AI Gemini, defended by two
layers (sandbox-exec containment plus human-in-the-loop approval), with no
analysis or GUI subsystems. It is drop-in compatible with the project
ecosystem: it reads an existing project's AGENTS.md / CLAUDE.md /
.mcp.json (and Claude Code-format skills) as-is, so one project setup
serves every runtime that works on it. It began life as a Claude Code
backup and was repositioned as an independent runtime once real-world use
outgrew that role
(ADR-0061).
brew install nlink-jp/tap/gem-agent
mkdir -p ~/.config/gem-agent
cp config.example.toml ~/.config/gem-agent/config.toml # set [gcp].project and [model].namecd /path/to/your/project
gem-agent # interactive REPL
gem-agent "run the tests" # interactive, with this as turn 1
gem-agent -c # continue the last session here
gem-agent sessions # list resumable sessions
gem-agent trust # this project's trust state and trusted files
gem-agent -p "summarize this repository" # one-shot, pipe-friendlyThe current directory becomes the project: file tools cannot leave it,
sandboxed shell commands cannot write outside it, and mutating tool
calls ask for approval before running. Each session also gets a work
directory of its own for anything that is not part of the project —
intermediate data, an oversized tool result, a screenshot a server
returned — so the working copy stays clean. $GEMAGENT_WORK_DIR names
this session's, and nothing in it is ever deleted for you: gem-agent workdirs lists what earlier sessions left behind, and workdirs clean
removes it after showing you exactly what and asking first. Requirements: macOS (Apple
Silicon), a Google Cloud project with Vertex AI enabled, and ADC
(gcloud auth application-default login) — details in
configuration.
Each line below is expanded in a focused document under docs/en/reference/; the ADRs hold the reasoning.
Interface — an inline Bubble Tea
TUI with native scrollback, a bottom-pinned input box that stays live
during a running turn (Enter queues the next message), a live turn
status — stream heartbeat, stall warning, visible retries, and the
model's thought summaries streaming as it thinks — a Ctrl+C that
always comes back (file walks stop within one syscall, anything else
is abandoned after a bounded grace — ADR-0065), IME-friendly
approval dialogs,
Tab completion for @-paths, /-commands, and skill
names, !command shell escape, mermaid fences drawn in place in the
reply — as pictures where the terminal draws images (flowchart / sequence /
ER / pie / state / gantt / mindmap, CJK labels included), as box art elsewhere
(flowchart / sequence / ER, CJK labels included, from the same engine; pie,
state, gantt and mindmap stay source there); anything that cannot be
drawn right stays source — fifteen slash commands (/help
/tools /mcp /auto /readonly /compact /settings /riskbook /usage
/memory /skills /skill /version /clear /quit), a provenance-first /settings
panel, theme control, and a fully bilingual chrome
([tui].language = auto|ja|en). A positional argument is the first
interactive turn — gem-agent "…" runs it and hands you the keyboard
(ADR-0064). Pipes fall back to a plain REPL;
-p runs one-shot (mutating tools denied, and the denial tells the
model what still runs rather than to ask you; --allow grants named
tools per run, --auto arms the risk ladder — ADR-0053), and
data | gem-agent -p "…" attaches piped stdin as isolated data,
never as prompt text (ADR-0055). The pipe is read to EOF; if it is
still open after 2 s, a stderr line says so and names the remedy —
launch with < /dev/null when nothing is meant to be attached
(ADR-0067).
Read-only is a mode — a session
ceiling on a separate axis from approval: --read-only caps the session
at the read lane, so a write, a memory save or a shell command asking
for more is refused before any gate rather than escalated to one, and
nothing outside the session scratch changes. --auto-read-only lets the
runtime turn it on when you ask for a session that changes nothing; it
only ever tightens. /readonly on|off and /readonly auto on|off move
either in a session, --writable opts a run out of a configured one,
and the status line and startup banner carry the state. A call over the
ceiling asks whether to lift the mode — a different question from
approving the call, which is asked separately afterwards (ADR-0080).
Built-in tools — orientation
(list_files/list_tree/search_files, ignore-aware: dependency and
build directories and .gitignore'd content are skipped with every
skip reported), reads windowed by line or by byte — the tail of a
saved single-line result included — and
summaries (read_file/summarize_file), delegated project search in
an isolated child context (agentic_file_search), atomic batched edits with
diagnosed misses (edit_file/write_file, with a shrink guard so a
whole-file rewrite cannot silently summarize a document away), file identification with
hashes (file_info), images and documents for the model
(view_image/read_document), an image drawn on YOUR screen
(show_image), a sandboxed shell whose read, write and
operator lanes the kernel enforces (shell_exec), a
deterministic clock/calendar (datetime), the model's own runtime
picture (agent_info), structured mid-turn choices (ask_user), and grounded web access
(web_search/web_fetch).
Attachments — @file,
@dir/, screenshots (@~/Desktop/…, @clipboard), PDFs and Office
documents, and audio/video — routed through your GCS bucket when
configured, inline otherwise.
Inline images — the model shows you a picture with show_image, and
you ask for one with /show <path> — PNG or JPEG, up to 2 MiB; the path may
hold spaces, be quoted, or be a file dragged into the window; it is drawn in the terminal when the
terminal can draw ([tui].images, auto by default; iTerm2 and kitty, off
inside tmux and screen). view_image is the other direction — that is the
model looking at an image, not you
(ADR-0089 /
ADR-0090 /
ADR-0091).
Approval and safety — per-call
MITL gates that show the model's own declared purpose for the call
alongside its arguments, a deny-with-reason answer (N) that carries
your one-line "do this instead" to the model inside the denial itself,
a session allowlist that never covers
Block-tier calls, an opt-in two-tier auto-approve (rules first, model
review second — edits to instruction and configuration files such as
AGENTS.md and .mcp.json always ask you, a read of a credential-named
file — .env, keys, credential stores — asks you in every mode and is
denied in -p, performed in a sandboxed child that cannot open
credential material at all, so the list raises the question and the
kernel is what refuses, while the walks list them and name what they
could not read,
and a trusted project's
instruction and configuration files are pinned by content so a changed
one asks again before it is loaded), shell commands judged by
the Seatbelt lane they declare rather than by their text (a read-lane
command runs unasked — compile, vet and test included, the toolchain
cache riding the session scratch in every lane; the write lane cannot
touch AGENTS.md or .git/config; the operator lane is yours alone —
ADR-0084), a runtime note when an
MCP server answers three calls in a row with the same error — the model
is told whose words the error is and to report to you rather than
investigate — a per-tool approval policy with scope-aware resolution and
trust-gated project loosening, a layered risk rulebook the auto-mode
reviewer reads — hand-written or drafted from your own recorded
answers, never skipping a gate — the Seatbelt sandbox, operator pre-tool
hooks that run the same guard scripts Claude Code does and refuse a
call before the ladder sees it, session-start and prompt-submit hooks
on the same contract whose output reaches the model as labelled data,
startup gates for
broad roots and first-seen projects, and nonce-tag isolation of all
tool output — which also keeps the request prefix byte-stable for
81–95% measured context-cache hits. Text the model writes, and the other
text from outside the runtime the TUI shows, reaches the terminal with its
control characters removed, so a prompt-injected reply cannot retitle the
window, write the clipboard, move the cursor or disguise a command in an
approval dialog; -p output stays byte-for-byte when it goes to a pipe or
a file.
Sessions — full-fidelity JSONL
transcripts, per-project state layout, --continue/--resume (UUID session ids, resumable by prefix) with
deliberate refusals (wrong directory, wrong model), automatic context
compaction with an honest notice, the /usage token statement, and
approval-gated agent memory across sessions.
Integration — drop-in reading
of AGENTS.md/AGENT.md/CLAUDE.md/GEMINI.md up the directory
tree, Claude Code-format .mcp.json MCP servers (global + project),
and Claude Code-format skills with progressive disclosure (a loaded
skill names its directory, as in Claude Code, so its own scripts run)
— both reloadable mid-session (/mcp reload, /skills reload), with
--mcp on|off to switch MCP per run. [mcp].exclude names the MCP
servers, or the single functions of them, that a session does not have:
an excluded server is never started, an excluded function is never
declared, and a read/write server can be kept for its read half (ADR-0077).
[mcp].advertise = "on-request" (an opt-in) keeps every server
connected but declares a tool to the model only once loaded — by the
model through find_tools, a librarian side call that reads the whole
catalogue and names the tools for the task, or mcp_load; or by you
through preload, --allow, or /mcp load — and withholds tools whose
descriptions address the assistant instead of describing a tool
(ADR-0083).
Configuration — the full config reference, precedence, CLI flags, content-filter behaviour, endpoint notes, and opt-in audit logging to Cloud Logging (default) or your OTLP collector (metadata only, never conversation content, and nothing added to startup).
Out of scope by design: RAG or vector memory, data analysis, GUI, non-macOS platforms.
gem-agent is built, signed and tested for macOS on Apple Silicon only,
and the protections described above rest on macOS's kernel sandbox
(sandbox-exec). At every start gem-agent measures that sandbox with
real probes; the result decides how commands are treated:
- verified — the startup banner shows no sandbox warning: read-lane commands run unasked, write-lane commands ask, and the kernel bounds both.
- unverified — the banner says
sandbox unverified: …orsandbox: DISABLED: every shell command asks you before it runs, and nothing else stands between the command and your machine. The banner says which, and/settingsshows the measured state on request.
A copy of the source rebuilt for another platform (Linux, WSL, Windows)
has no kernel sandbox behind it. If it starts without the warning above,
its sandbox check was altered and its behaviour is not the one this
README describes — treat it as unverified whatever it prints. Such builds
are not supported here; the tests under internal/sandbox are the bar
any port has to meet.
make build # outputs dist/gem-agent
make testdocs/en/INDEX.md is the entry point (日本語) — one catalog rather than a list maintained in several places. It covers the RFP (the canonical spec), the seven feature references linked above, architecture, the monthly drill, promotion criteria, and every ADR.
MIT