Skip to content

Latest commit

 

History

28,193 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OpenAgents

OpenAgents builds agent infrastructure: typed decisions, bounded execution, programs, permissions, traces, and evaluation. Coder is the coding product, with terminal and headless interfaces and native mobile readers for saved Codex and Claude Code chats. The suite plan extends the same runtime to mobile, web, cloud execution, and computer control. Coder One supplies the configurable agent components used in Terminal-Bench experiments and Coder's delegate execution. Microcoder is a separate experimental loop that combines Jev, generation, and a shared knowledge base. Gym measures results and lets you inspect and replay the evidence.

The repository also contains decision-model implementations and services, Rust SDKs and CLIs, a Nostr relay, public protocol specifications, and Voyager's Minecraft agent. Rust Native supplies the experimental shared UI foundation: typed views and generic styles. Coder's application palette lives separately in coder-ui, used through the terminal's compatibility exports. The mobile readers render Rust Native lists and transcripts through thin SwiftUI and Android framework controls. Its Verse home screen mounts the shared desktop world through Rust Native's generic drawing-surface contract and native Metal (iOS) or GLES (Android) surfaces. Product state and transport are Rust; native glue also includes the Swift bridge for Apple's on-device model. Python and shell handle training, benchmark acquisition, and infrastructure.

The general agent architecture and optimization design describe the broader integration target. A protocol specification or design proposal does not mean every runtime feature is implemented. The glossary labels implemented, partial, and proposed concepts.

Agent labor is a high-priority development track: let independent operators offer bounded coding work and receive Bitcoin for accepted results. The labor market plan and Coder network plan connect this work to reusable knowledge, programs, and measured outcomes.

The OpenAgents protocol index contains 27 authored NIPs plus shared contracts. Encrypted artifacts, free market negotiation, and labor-term validation now have Rust components. A recoverable free labor host links an exact order to bounded execution and acceptance. Paid settlement and the complete multi-operator market remain unfinished. The implementation coverage report maps every contract to its implemented parts and remaining work.

NIP-SOV restores the historical sovereign-agent design as a Designed profile for durable identity, bounded initiative, custody, guardians, and market participation. It composes the current NIPs; custody, sovereign lifecycle, and treasury adapters remain to be built.

Coder iOS 0.5.0 (50) is available in internal TestFlight. The plaza has two zone portals. Ruins (formerly the Atlantis forest) runs the original real-time Wizard Woods combat with the Firebolt, Magic Missile, and Fireball hotbar. Lagrange 1 is a construction station orbiting the Sun–Earth L1 point, with restricted three-body orbital mechanics, station-keeping, true-size Sun, Earth, and Moon, and a rocket-equation maneuvering pack: grab parts at the depot and latch them into a keel jig. Use Map → Ruins portal or L1 portal to enter and Plaza to return. Build and register new zones. The build 50 record has the release receipt and simulator evidence. Physical-device acceptance is separate.

New or unconfigured installs still join wss://relay.openagents.com automatically for shared plaza presence. Custom relays and an explicit Leave choice persist. Tap the physical GYM board to open its world-rendered entry; detailed Gym records require a separate host grant. The build 48 connection and Gym evidence and signed relay receipts remain available. Shared zone state, saved L1 assemblies, and creator publishing remain roadmap work.

Computer setup uses ./pair or installed coder pair. Chats load after pairing, the world relay choice persists, and movement combines with looking or double-tap jumping. Pinch to zoom; the crosshair recenters the camera. See the connection corrections.

Verse fills the screen, including behind the system clock. Motion look follows body turns and upward tilt with interpolated camera movement; hold the left side to walk. Touch look remains available. The computer now renders its prompt on the physical 3D monitor; approach it and tap its screen. The world computer pairs by QR code or a pasted string and opens saved Codex and Claude transcripts with follow updates and encrypted local caching. Pair your computer. The Android app uses the same Rust reader and Verse runtime. Its emulator acceptance is recorded separately from TestFlight and physical-device release acceptance. The Gym building loads Microcoder and Terminal-Bench boards only while you are inside. Inspect recorded charts and explicitly request host-enabled runs through a separate Gym connection. Motion-camera and build 44 verification, world-computer verification, and full-screen release evidence.

Start here

Goal Guide
Find documentation and the complete direction Documentation index, master roadmap, launch roadmap, playtesting program, catalog, glossary
Run the coding agent Install Coder, headless mode
Track suite implementation and next work Migration status and issue map, local task commands, execution owner and evidence
Control an existing task over Nostr Scoped host/client bridge: explicit pairing, observe/steer/cancel rights, retained retries, and bounded private history
Install a bounded task host Verified bundles, one-shot services, and rollback, platform acceptance and limits
Try the experimental knowledge-assisted loop Microcoder, shared knowledge base
Inspect exact knowledge inputs and comparisons Private and immutable bundles, evidence integrity, frozen study bookkeeping
Read computer chats on your phone Pair the Coder mobile reader, iOS build and TestFlight setup, Android build and emulator, verification
Review mobile platform feasibility Rust native prototype and measured limits
Build shared terminal, web, and native UI Rust Native, framework contract, Coder architecture, build order, adoption map
Compare saved agent transcripts Gym head-to-head replay
Inspect benchmark results Terminal-Bench status and evidence
Run a benchmark or controlled experiment Harness runbook, experiment template
Use typed decisions Decision models, caller CLI, Rust clients
Run decision services Gateway, deployment
Operate the Nostr relay Local relay runbook, production (Cloud Run) runbook, configuration
Review protocol support and gaps NIP implementation coverage, implementation plan
Operate a Minecraft agent Voyager, watch an episode
Walk the Verse world on desktop or your phone Verse, mobile controls, maps, companions, and gates, loaded zones: Ruins and Lagrange 1, Gym building and run boards

Run Coder

Use the toolchain pinned in rust-toolchain.toml, currently Rust 1.97.1. From the repository root:

cargo run -p coder --bin coder
cargo run -p coder --bin coder -- -p "count the crates"
cargo run -p coder --bin coder -- -p --json --trace one.jsonl "count the crates"

To install this repository's build on your PATH:

./scripts/install-coder.sh
coder --version
coder doctor

The installer records the build it replaces; ./scripts/install-coder.sh --rollback restores that build. coder --version identifies the repository and commit. coder doctor reports the selected execution backend, credentials found, boundary support, and trace location without running a turn. See the installation guide.

The default CODER_DELEGATE=auto uses Microluna in process when the Codex login has more than ten minutes left on its access token. Otherwise it tries authenticated Claude Code, then Codex CLI, then the Open Responses door. Coder One supplies the probes, briefing, and configurable execution loop. Commands run inside the host's filesystem boundary. CLI follow-ups resume their session; Microluna receives the earlier conversation as context. Terminal and headless modes use the same turn implementation.

Setting Purpose
CODER_DELEGATE=auto|always|off Use an available executor, require one, or disable this delegation path.
CODER_DELEGATE_AGENT=microluna|claude-code|codex Select the executor.
CODER_DELEGATE_MODEL Select its model.
TYPESAFE_API_KEY or ~/.openagents/jev.json Configure Jev for the delegate briefing. Without a key, that briefing carries the request alone.
CODER_DOOR_KEY, CODER_DOOR_URL, CODER_MODEL Configure the Open Responses fallback.
CODER_WORKER, CODER_RELAY Explicitly select a relay worker.
CODER_SHELL=off Disable command execution; a delegated turn is read-only.

An explicitly selected relay worker or local executor takes precedence over automatic delegation. With no configured executor or generation credentials, the fallback is a labeled stub response. Use coder doctor to establish what will actually run. The delegate guide documents selection, credentials, boundaries, usage, and session continuity.

For development from another project's directory:

alias coderdev=~/work/openagents/scripts/coderdev
coderdev -p "count the crates"

Adjust the alias to your checkout. The launcher builds before running, keeps the caller's working directory, forwards arguments, and prints the revision and binary it launched. CODERDEV_ENV_FILE selects a private environment file; CARGO_TARGET_DIR selects the build directory. Use a separate target directory for each worktree.

Coder records conversations as ATIF traces in ~/.openagents/traces/ by default. Delegated turns also retain their briefings and native executor streams. See traces and headless output and exit codes.

Inspect and replay runs in Gym

Open the Terminal-Bench Runs view:

cargo run -p gym --features tui --bin gym-terminal -- --terminal-bench

Press Enter for a run summary, t for its transcript, or p for head-to-head replay. Choose a task and one attempt on each side. Coder One versions and repeated attempts remain separate choices. You can also compare two local attempts or view a public attempt on its own.

Press w for an experiment pulse: per-arm results, uncertainty, component outcomes, check calibration, and the stopping verdict. The same view is available as gym experiment pulse ID; add --jev for cached judgments or --live for advisory assessments of running trials. Unknown trial costs stay explicit, and restarting an experiment preserves its stopping policy. See the experiment guide and September 24 issue review.

Press l in the head-to-head picker to switch between newest first and Jev's learning order. Both Coder and Fable attempts receive the same learning judgments used in Runs. The picker shows scores and reasons; the selected task and attempts stay selected as scores arrive. During replay, l pauses the clock and shows both full Jev assessments; press it again to return to the transcripts at the same point.

New analysis uses TYPESAFE_API_KEY or ~/.openagents/jev.json and shows progress and estimated cost. Answers are cached. Add --no-jev to disable new calls while keeping cached assessments available. Transcript playback itself makes no model calls. See learning from comparisons for scoring, cache behavior, and evidence limits.

Replay key Action
Space Play or pause both transcripts.
l Switch between chronological replay and both runs' Jev assessments.
+ / - Change speed through 1×, 2×, 5×, and 10×.
Left/right arrows Seek backward/forward 30 seconds.
n / b Jump to the next/previous event.
End Reveal both complete transcripts.
Tab, then arrows or Page Up/Down Choose a side and scroll it independently.
d Switch between readable conversation and full records.
Escape Return to the task and attempt picker.

Playback starts paused. 0 / N events means the transcript is loaded but the clock has not reached its first event. Press Space, n, or End. The two sides share an elapsed-time clock aligned to their own starts. Recorded and estimated timestamps are labeled; source step timestamps do not imply token-by-token streaming. Replay reads saved evidence and does not rerun the agents.

Gym uses the shared Coder Markdown renderer for messages and reasoning in both transcript views, including headings, emphasis, lists, quotes, links, tables, and fenced code. Commands, tool output, and d full records remain literal. Long code and output stay scrollable without dropping lines.

Click an underlined file path in either transcript view to inspect its retained contents. The file viewer wraps, scrolls, labels historical reads and later snapshots, and returns to the transcript with Esc. Opening it pauses head-to-head playback.

Load the public transcripts on each computer

The pinned Fable 5.1 collection contains 1,650 listed attempts across 66 TB4 tasks, with 1,649 published transcripts across five effort settings. One attempt has no published trajectory. The transcript bodies occupy about 6 GB and are stored outside Git at ~/.openagents/terminal-bench/public-replays/.

Pulling this repository downloads the attempt catalog, not those public transcript files. On each computer where you want to replay them, use uv to load the pinned collection:

(cd bench/terminal-bench && uv run python -m tbench.public_replays)

The downloader resumes and verifies files against their retained SHA-256 digests. A [not on this computer] entry needs its local file; the failed pane shows the cause and recovery instructions. After acquisition, press Escape and Enter to reload the pair.

Local attempts come from ~/.openagents/terminal-bench/jobs/ and the repository's retained traces. To copy available transcripts from a separate benchmark host over SSH:

(cd bench/terminal-bench && uv run python -m tbench.sync_replays HOST)

Replace HOST with its SSH name or address. The mirror lives under ~/.openagents/terminal-bench/replay-jobs/; it is a snapshot of retained records, not a live stream. Missing or incomplete evidence stays explicit.

See the full replay guide, Gym terminal controls, and Gym CLI. To start directly in replay, add --head-to-head to the Gym command above. --print provides a noninteractive view. The plain gym-terminal command without --terminal-bench opens the decision-model views.

Coding-agent benchmark evidence

The latest Microcoder development results show knowledge-assisted passes on three selected tasks. The strongest individual efficiency result is a fin-saccr-rwa pass in 2:48 for $0.0404, against Fable 5.1 low's three public passes at 3:42–4:28 and $1.23–$1.49. This is an in-sample development result with knowledge learned from the task; it does not establish general superiority or the effect of adding Coder to an otherwise identical configuration.

Microcoder figures are per-run Luna, Jev, and embedding costs. They do not include the cost of developing the knowledge base.

Task Reported development result Efficiency observation
fin-saccr-rwa 4/4 with SA-CCR knowledge entry v9 The 2:48 pass was faster than all three Fable low passes and cost about 1/30 of its cheapest recorded pass. Two of the four runs share a mixed record directory.
embedding-drift-monitor 8/9 with knowledge A 2:21 pass cost $0.0165 and was faster than four of Fable low's five passes; the median successful Microcoder run was slower.
gsea-proteomics 4/4 with relay-supplied knowledge Every pass cost less than Fable low's cheapest pass. The decisive entry came from a winning Fable trace on this task.

The knowledge-base guide covers retrieval, harvesting, admission, and NIP-KB sharing. GSEA and SA-CCR runs received their entries from a Nostr relay. Their knowledge was developed using these tasks and public winning traces, so those tasks cannot establish generalization. The results ledger reports no out-of-sample Microcoder passes. Earlier failures remain recorded, and interrupted or credit-exhausted runs are ungraded rather than counted as successes or failures.

Microluna v19-fire also passed embedding 5/5 at about $0.0153 per run, excluding the fire-loop judge's Jev cost. Its 5:26 median was slower than Fable low's 2:55. This task was used to develop the harness and its method check. See the fire-loop guide for the development workflow.

The retained same-executor controller experiment is a separate result. It holds Claude Code, Opus 5.5, medium effort, tools, system prompt, and outer budgets fixed across ten TB4 tasks, with three attempts per task per arm:

Arm Passes Total model usage cost Mean agent minutes per attempt
Plain Claude Code 15/30 $27.04 6.6
Coder One v8 controller 18/30 $45.42 14.6

The controller cost 68% more and took 2.2× the agent time. The three extra passes were not statistically established as a general gain (exact McNemar p = 0.51). Persistence produced the specific mvcc-lsm-compaction win, 3/3 versus 0/3, and most of the additional cost. Selected Microcoder wins do not overturn that controlled result.

Negative studies remain available. The v18 family passed 0/9 confirmation and 0/9 development attempts; setup changes and a restart left its strict protocol result inconclusive. The 72-candidate truthful-checks confirmation missed its declared joint-improvement requirement. The later 90-attempt protocol and launch records establish the frozen plan and recorded launch; they do not establish its current status or a completed measurement.

Use the results ledger for per-task comparisons, the report index for historical studies and retained traces, and the data-quality notes for accounting and grading limitations. These are repository snapshots, not the execution host's live queue. The tunable policy guide explains the controller components; a policy's presence in code is not a measured result.

Implementation map

Crates Responsibility
coder Terminal and headless turns, typed routing, delegation, the shell loop, and the coder-worker relay client.
coder-one Standalone issue-to-PR agent and reusable components for probing, briefing, execution, checks, repair, escalation, and persistence.
microluna Short model sessions with five native tools, host-enforced execution, and ATIF traces; used by Coder One and Coder's delegate path.
microcoder, knowledge Experimental Jev-guided coding loop, knowledge retrieval and expansion, entry admission, and NIP-KB publication and synchronization.
coder-project, coder-scheduler Supervised project programs, deterministic task admission, durable scheduling records, and simulations.
rust-native Experimental semantic views, typed intents, generic style composition; application palettes and native adapters are separate.
coder-ui Coder application palette and presentation values; separate from the reusable UI framework.
coder-terminal Terminal design system, composer, frames, and rendering; re-exports Coder's coder-ui theme for existing consumers.
coder-boundary, supervise Filesystem enforcement, workspace snapshots, process-group cleanup, deadlines, and output bounds.
atif, receipts Append-only agent trajectories and versioned execution receipts.
capability Capability manifests, host-owned trust, and bounded probes.
coderbench Whole-episode task manifests, workspace checks, and recorded goldens.
gym Decision-model suites, result chains and gates, experiment analysis, run inspection, and transcript replay.
jev, oak Typed Rust clients, the decision API CLI, and MCP caller tools.
kev, laya, lev Local decision-model implementations; Lev uses the Swift bridge to Apple's FoundationModels.
gateway, tenancy Authenticated HTTP serving, artifact admission, durable quotas, accounts, billing, skills, and training records.
discovery Shared documentation, agent cards, skill indexes, and discovery surfaces.
plugin, plugin-pdk, plugin-outline Bounded plugin host, shared packet ABI, and diagnostic guest.
plugin-repo-map, plugin-code-search, plugin-test-report Evidence guests that programs/evidence-guests.json runs; built by scripts/build-plugin-guests.sh.
nostr, nostr-relay Protocol verification and the self-hostable PostgreSQL-backed relay.
voyager Minecraft curriculum, bounded programs, critics, skill retention, and traces through the separate nightly Rust bridge.
gym-bridge Private Gym observation and explicit recipe launches over Nostr; portable client plus a separately enabled local host.
verse Shared desktop/iOS/Android world, collision-aware map navigation, companion reactions, local item-dependent gates, native GPU surfaces, and multiplayer presence over NIP-MV. Local gate choices do not grant service access or synchronize inventory.

The host owns permissions, deadlines, budgets, and execution boundaries. Typed judgments inform decisions; their shape does not establish correctness or grant authority. The generation interface supports multiple backends, and workload-specific evaluation determines whether a replacement helps.

Architecture and protocols

The NIP directory has three lanes: 99 official specifications, 17 Block/Buzz extensions, and 27 OpenAgents NIPs plus shared contracts. nips/manifest.json pins the upstream revisions; the September 26 review records their changes. A source inventory does not establish complete implementation.

Current implementation status, September 26, 2026:

Area Implemented scope Remaining boundary
Official protocol updates Petnames, comments, highlights, emoji, authentication hints, relay-access declarations, payment-target parsing, and atomic allow/ban updates. Client helpers have specific roles; parsing a declaration does not run a membership or payment service. See the official ledger.
Private relay data Author-only NIP-78 app state, author/recipient visibility for encrypted artifacts, and exclusion of private content from search. Hosts still authorize the actions described by those artifacts.
Block read-state snapshots Configured, authenticated HTTP snapshots from the writer database, with signature/digest checks, replay protection, and refusal of incomplete cuts. Client merge behavior and cross-subscription synchronization remain separate work. See Block support.
Other Block helpers Persona adoption checks, thread-batch parsing and bounds, and federated-identity policy checks with a required external verifier. Complete launchers, thread query service, JWT/JWKS integration, push delivery, and managed-agent lifecycle remain unfinished.
Agent markets and labor Free labor host: authenticated agreement, explicit bounded execution, retained delivery, separate buyer checks and acceptance, duplicate/restart recovery. The local synthetic acceptance uses distinct keys under one operator. Paid settlement, resolver execution, nonzero rework, discovery, and production service remain unsupported.
x402 Lightning Offline BOLT11 signature, amount, payee, expiry, request-binding, and preimage validation for HTTP/MCP and the explicitly selected native Nostr profile. Wallet authority, durable proof consumption, execution recovery, and live payment interoperability remain unfinished.

Relay configuration changed: incomplete NIP-PL push configuration now fails at startup. Reserved PMA events and unsupported CW thread query modes are refused; unsupported roles are not advertised. See the configuration contract before upgrading a relay configured for those paths.

The coverage report lists each contract's implementation and gaps. The implementation plan defines completion evidence, beginning with a durable free-labor rehearsal. Session, workspace, tracked-work, automation, environment, and live-media profiles also retain substantial runtime work. No complete-NIP or operational-market claim follows from the new validators.

The teardown integration plan maps all 81 archived teardown documents into Coder. Six new draft profiles cover sessions, workspaces, tracked work, automation, environments, and live media, with governed preferences and component updates in existing contracts. The coverage ledger links every source and separates specifications from implementation work.

The x402 Lightning integration plan and draft NIP-X402 specify paid operations with Nostr discovery and private evidence, standard HTTP/MCP compatibility, and an opt-in native Nostr profile. Offline invoice and binding validation now exists; wallet, settlement, and recovery services remain pending. Upfront tool purchases are separate from labor payment after acceptance; zaps are not substituted for x402 proofs.

The implemented Coder relay path carries signed, NIP-44-encrypted NIP-CJ jobs between the terminal and a worker. NIP-42 authenticates relay connections. The relay transports ephemeral job events; the worker holds provider credentials. This is one execution path alongside local delegates and HTTP backends. The HTTP gateway separately uses bearer-key admission. See the relay measurement and local relay runbook.

For further design and operation:

Verify and contribute

For daily Rust work, run targeted checks on the pinned toolchain:

./scripts/verify-rust.sh --crates coder

A bare invocation selects changed packages. The full matrix is release-only: ./scripts/verify-rust.sh --release. It never blocks ordinary issue work, commits, or pushes. Direct focused Cargo checks are also valid.

Read verification.md for scope, feature coverage, external prerequisites, and optional checks. Documentation-only changes require link, path, and artifact checks rather than the Rust gate. Required checks run on contributor machines or non-GitHub infrastructure; this repository does not use GitHub-billed automation.

The September 26 verification record covers the latest protocol implementation and its two test-fixture repairs. It records 290 passing Nostr library tests, passing default and feature-enabled workspace tests and strict Clippy, dependency checks, and live PostgreSQL acceptance, including snapshots, privacy, restart, backup/restore, and actual Coder/worker processes. Gym's feature suite passes 595 tests, with one ignored.

That historical protocol record combines a full run with scoped recoveries: the original full run remains marked failed, and the successful scoped runs remain marked partial. The records retain the earlier failures, exact code revisions, and fixes for host-dependent Gym metadata and webhook test synchronization. Metal, long-running soak, external model and wallet integration, and production deployment are outside that evidence.

The replay-rendering record and experiment safeguards review retain earlier feature-specific results. Documentation-only README updates check links and formatting without rerunning the Rust suite.

AGENTS.md is the contributor contract. See LICENSE and the dependency and provenance policy for the repository's licensing records and dependency requirements.

About

Monorepo & docs

Resources

Stars

453 stars

Watchers

8 watching

Forks

Used by

Contributors

Languages