Skip to content
nlink-jpPublic

About

Interactive CLI agent runtime on Vertex AI Gemini — a deliberately minimal, auditable loop of read, edit, shell, MCP and approval, defended by sandbox-exec containment plus human-in-the-loop approval, with no analysis or GUI subsystems. Drop-in with existing projects: it reads their AGENTS.md, CLAUDE.md and .mcp.json as they are

Topics

Resources

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

gem-agent

Interactive CLI agent runtime backed by Vertex AI Gemini 3.x — deliberately minimal, sandboxed, and usable both interactively and headlessly.

Released. Install with brew install nlink-jp/tap/gem-agent or from the releases page (Developer ID signed, Apple-notarized, macOS arm64) — the releases page is the authority on the current version, so this line does not repeat it. It is a cli-series tool: interface stability is a promise, and breaking changes go through the org's breaking-change process (ADR-0061). See the RFP for the full specification.

日本語版 README

Why

gem-agent is an independent agent runtime: a minimal, auditable agent loop (read / edit / shell / MCP / approval) on Vertex AI Gemini, defended by two layers (sandbox-exec containment plus human-in-the-loop approval), with no analysis or GUI subsystems. It is drop-in compatible with the project ecosystem: it reads an existing project's AGENTS.md / CLAUDE.md / .mcp.json (and Claude Code-format skills) as-is, so one project setup serves every runtime that works on it. It began life as a Claude Code backup and was repositioned as an independent runtime once real-world use outgrew that role (ADR-0061).

Quickstart

brew install nlink-jp/tap/gem-agent
mkdir -p ~/.config/gem-agent
cp config.example.toml ~/.config/gem-agent/config.toml   # set [gcp].project and [model].name
cd /path/to/your/project
gem-agent                                  # interactive REPL
gem-agent "run the tests"                  # interactive, with this as turn 1
gem-agent -c                               # continue the last session here
gem-agent sessions                         # list resumable sessions
gem-agent trust                            # this project's trust state and trusted files
gem-agent -p "summarize this repository"   # one-shot, pipe-friendly

The current directory becomes the project: file tools cannot leave it, sandboxed shell commands cannot write outside it, and mutating tool calls ask for approval before running. Each session also gets a work directory of its own for anything that is not part of the project — intermediate data, an oversized tool result, a screenshot a server returned — so the working copy stays clean. $GEMAGENT_WORK_DIR names this session's, and nothing in it is ever deleted for you: gem-agent workdirs lists what earlier sessions left behind, and workdirs clean removes it after showing you exactly what and asking first. Requirements: macOS (Apple Silicon), a Google Cloud project with Vertex AI enabled, and ADC (gcloud auth application-default login) — details in configuration.

What it does

Each line below is expanded in a focused document under docs/en/reference/; the ADRs hold the reasoning.

Interface — an inline Bubble Tea TUI with native scrollback, a bottom-pinned input box that stays live during a running turn (Enter queues the next message), a live turn status — stream heartbeat, stall warning, visible retries, and the model's thought summaries streaming as it thinks — a Ctrl+C that always comes back (file walks stop within one syscall, anything else is abandoned after a bounded grace — ADR-0065), IME-friendly approval dialogs, Tab completion for @-paths, /-commands, and skill names, !command shell escape, mermaid fences drawn in place in the reply — as pictures where the terminal draws images (flowchart / sequence / ER / pie / state / gantt / mindmap, CJK labels included), as box art elsewhere (flowchart / sequence / ER, CJK labels included, from the same engine; pie, state, gantt and mindmap stay source there); anything that cannot be drawn right stays source — fifteen slash commands (/help /tools /mcp /auto /readonly /compact /settings /riskbook /usage /memory /skills /skill /version /clear /quit), a provenance-first /settings panel, theme control, and a fully bilingual chrome ([tui].language = auto|ja|en). A positional argument is the first interactive turn — gem-agent "…" runs it and hands you the keyboard (ADR-0064). Pipes fall back to a plain REPL; -p runs one-shot (mutating tools denied, and the denial tells the model what still runs rather than to ask you; --allow grants named tools per run, --auto arms the risk ladder — ADR-0053), and data | gem-agent -p "…" attaches piped stdin as isolated data, never as prompt text (ADR-0055). The pipe is read to EOF; if it is still open after 2 s, a stderr line says so and names the remedy — launch with < /dev/null when nothing is meant to be attached (ADR-0067).

Read-only is a mode — a session ceiling on a separate axis from approval: --read-only caps the session at the read lane, so a write, a memory save or a shell command asking for more is refused before any gate rather than escalated to one, and nothing outside the session scratch changes. --auto-read-only lets the runtime turn it on when you ask for a session that changes nothing; it only ever tightens. /readonly on|off and /readonly auto on|off move either in a session, --writable opts a run out of a configured one, and the status line and startup banner carry the state. A call over the ceiling asks whether to lift the mode — a different question from approving the call, which is asked separately afterwards (ADR-0080).

Built-in tools — orientation (list_files/list_tree/search_files, ignore-aware: dependency and build directories and .gitignore'd content are skipped with every skip reported), reads windowed by line or by byte — the tail of a saved single-line result included — and summaries (read_file/summarize_file), delegated project search in an isolated child context (agentic_file_search), atomic batched edits with diagnosed misses (edit_file/write_file, with a shrink guard so a whole-file rewrite cannot silently summarize a document away), file identification with hashes (file_info), images and documents for the model (view_image/read_document), an image drawn on YOUR screen (show_image), a sandboxed shell whose read, write and operator lanes the kernel enforces (shell_exec), a deterministic clock/calendar (datetime), the model's own runtime picture (agent_info), structured mid-turn choices (ask_user), and grounded web access (web_search/web_fetch).

Attachments — @file, @dir/, screenshots (@~/Desktop/…, @clipboard), PDFs and Office documents, and audio/video — routed through your GCS bucket when configured, inline otherwise.

Inline images — the model shows you a picture with show_image, and you ask for one with /show <path> — PNG or JPEG, up to 2 MiB; the path may hold spaces, be quoted, or be a file dragged into the window; it is drawn in the terminal when the terminal can draw ([tui].images, auto by default; iTerm2 and kitty, off inside tmux and screen). view_image is the other direction — that is the model looking at an image, not you (ADR-0089 / ADR-0090 / ADR-0091).

Approval and safety — per-call MITL gates that show the model's own declared purpose for the call alongside its arguments, a deny-with-reason answer (N) that carries your one-line "do this instead" to the model inside the denial itself, a session allowlist that never covers Block-tier calls, an opt-in two-tier auto-approve (rules first, model review second — edits to instruction and configuration files such as AGENTS.md and .mcp.json always ask you, a read of a credential-named file — .env, keys, credential stores — asks you in every mode and is denied in -p, performed in a sandboxed child that cannot open credential material at all, so the list raises the question and the kernel is what refuses, while the walks list them and name what they could not read, and a trusted project's instruction and configuration files are pinned by content so a changed one asks again before it is loaded), shell commands judged by the Seatbelt lane they declare rather than by their text (a read-lane command runs unasked — compile, vet and test included, the toolchain cache riding the session scratch in every lane; the write lane cannot touch AGENTS.md or .git/config; the operator lane is yours alone — ADR-0084), a runtime note when an MCP server answers three calls in a row with the same error — the model is told whose words the error is and to report to you rather than investigate — a per-tool approval policy with scope-aware resolution and trust-gated project loosening, a layered risk rulebook the auto-mode reviewer reads — hand-written or drafted from your own recorded answers, never skipping a gate — the Seatbelt sandbox, operator pre-tool hooks that run the same guard scripts Claude Code does and refuse a call before the ladder sees it, session-start and prompt-submit hooks on the same contract whose output reaches the model as labelled data, startup gates for broad roots and first-seen projects, and nonce-tag isolation of all tool output — which also keeps the request prefix byte-stable for 81–95% measured context-cache hits. Text the model writes, and the other text from outside the runtime the TUI shows, reaches the terminal with its control characters removed, so a prompt-injected reply cannot retitle the window, write the clipboard, move the cursor or disguise a command in an approval dialog; -p output stays byte-for-byte when it goes to a pipe or a file.

Sessions — full-fidelity JSONL transcripts, per-project state layout, --continue/--resume (UUID session ids, resumable by prefix) with deliberate refusals (wrong directory, wrong model), automatic context compaction with an honest notice, the /usage token statement, and approval-gated agent memory across sessions.

Integration — drop-in reading of AGENTS.md/AGENT.md/CLAUDE.md/GEMINI.md up the directory tree, Claude Code-format .mcp.json MCP servers (global + project), and Claude Code-format skills with progressive disclosure (a loaded skill names its directory, as in Claude Code, so its own scripts run) — both reloadable mid-session (/mcp reload, /skills reload), with --mcp on|off to switch MCP per run. [mcp].exclude names the MCP servers, or the single functions of them, that a session does not have: an excluded server is never started, an excluded function is never declared, and a read/write server can be kept for its read half (ADR-0077). [mcp].advertise = "on-request" (an opt-in) keeps every server connected but declares a tool to the model only once loaded — by the model through find_tools, a librarian side call that reads the whole catalogue and names the tools for the task, or mcp_load; or by you through preload, --allow, or /mcp load — and withholds tools whose descriptions address the assistant instead of describing a tool (ADR-0083).

Configuration — the full config reference, precedence, CLI flags, content-filter behaviour, endpoint notes, and opt-in audit logging to Cloud Logging (default) or your OTLP collector (metadata only, never conversation content, and nothing added to startup).

Out of scope by design: RAG or vector memory, data analysis, GUI, non-macOS platforms.

Supported platform

gem-agent is built, signed and tested for macOS on Apple Silicon only, and the protections described above rest on macOS's kernel sandbox (sandbox-exec). At every start gem-agent measures that sandbox with real probes; the result decides how commands are treated:

  • verified — the startup banner shows no sandbox warning: read-lane commands run unasked, write-lane commands ask, and the kernel bounds both.
  • unverified — the banner says sandbox unverified: … or sandbox: DISABLED: every shell command asks you before it runs, and nothing else stands between the command and your machine. The banner says which, and /settings shows the measured state on request.

A copy of the source rebuilt for another platform (Linux, WSL, Windows) has no kernel sandbox behind it. If it starts without the warning above, its sandbox check was altered and its behaviour is not the one this README describes — treat it as unverified whatever it prints. Such builds are not supported here; the tests under internal/sandbox are the bar any port has to meet.

Build

make build    # outputs dist/gem-agent
make test

Documentation

docs/en/INDEX.md is the entry point (日本語) — one catalog rather than a list maintained in several places. It covers the RFP (the canonical spec), the seven feature references linked above, architecture, the monthly drill, promotion criteria, and every ADR.

License

MIT

About

Interactive CLI agent runtime on Vertex AI Gemini — a deliberately minimal, auditable loop of read, edit, shell, MCP and approval, defended by sandbox-exec containment plus human-in-the-loop approval, with no analysis or GUI subsystems. Drop-in with existing projects: it reads their AGENTS.md, CLAUDE.md and .mcp.json as they are

Topics

Resources

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Releases

Contributors

Languages