Skip to content

Latest commit

 

History

293 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SpecPi logo

SpecPi

A Pi harness setup, built to specification.

npm version Build status MIT license

Website · Documentation · Evaluations · Releases

SpecPi Chat in VS Code: an open file beside the chat panel discussing a focused change.

SpecPi Chat · Example workspace


SpecPi 0.37.0 is a small starting point for the Pi coding agent. It is one opinionated setup for how the agent should work, not a marketplace of plugins.

At the center are two built-in extensions. Scope control keeps each task to the files it said it would touch. The improvement loop turns repeated friction into small, tested changes to the setup, instead of letting prompts and workarounds pile up. Around those are eight hand-picked packages, each locked to an exact version and checked before anything installs, plus SpecPi Chat, a VS Code panel for working alongside the agent.

It focuses on five things:

  • Control: clear scope, tool permissions, and lifecycle commands that ask before they change anything
  • Accuracy: exact version pins, checksums on state, rollback on failure, and proof over promises
  • Improvement: local notes become small, checked changes through /harness-improvement
  • Efficiency: delegation, persistent goals, and browser QA handled by the right tool for the job
  • Keep talking: long evals, builds and servers can run as background jobs, so the conversation is not stuck until they finish. The agent hears back when each one ends, and /jobs lists or stops them
  • Lean default: web access, browser QA, and delegation stay off until you need them. Turn them on for a session with /webaccess on, /browser on, and /delegate on — or let the agent ask when it hits the need, and answer the prompt

Everything it touches is written down, versioned, and easy to undo.

A local command guard

Before the agent runs a shell command, LANCET Nano, a 111M-parameter classifier, can check it on your own CPU. It reads Bash, PowerShell and cmd, takes about 10 ms for a typical command, needs no API key, sends nothing over the network, and costs nothing per check. Fixed rules settle the obvious cases first, blocking rm -rf / and passing git status, and the model scores the rest. A risky or uncertain verdict asks you before anything runs; /lancet-guard mode block blocks risky ones outright instead.

Bar chart of Triage Score (higher is better) for 14 command guards on the same 7,911 commands. LANCET Nano v0.4.3 scores 68.3, catching 77.0% of risky commands while stopping 9.0% of safe ones. Next are Jev 38.9, verdict-shell-safety 32.4, Kestrel 30.7, ModernBERT bash 23.2, dcg 19.8, AutoShell-0.8B 19.2, Gyra 17.9, bash-classify 14.3, Laya 13.3, bev-decider 13.3, laya-cli-gate 12.5, sh-guard 10.9 and Shieldstral-1.0-3B 6.1.

Every guard in the chart was scored once, at its own shipped setting, on the same three benchmarks, 7,911 commands in total, weighted by size: lancet-bench-2-next (1,602 risky commands and their safe look-alikes), the ShellRisk-Bench test split, and a 513-command neutral set. The Triage Score gives a guard a point for each risky command it asks about or blocks, and shrinks in proportion once it stops more than 10% of safe commands. The guard ships off. Run /lancet-guard setup in Pi once to download the model (about 109 MB, verified by SHA-256) and switch it on. How it works, what it costs, and its limits.

Measured context

This chart shows first-call context from a clean install: all eight pinned packages, the working agreement, and the skills Pi finds, in an interactive session with wishlist collection undecided. A headless session, or one with collection off, also leaves out the gap report and capability request tools. "Enabled" means browser QA, delegation, and web access are switched on, with no goal, scope, or improvement selection active.

The solid rows are measured by us, from the request each setup actually sends through one local test provider. That includes OpenCode, the DeepSeek Harness, and Oh My Pi, all measured as installed. The faded Codex CLI and Claude Code rows come from HarnessTax's published numbers, measured under their own setup. Treat those as a rough reference, not a head-to-head test. These are character counts. They say nothing about tokens, cost, or how well each tool does the job. The research page breaks down the enabled setup by feature, so you can see what each switch costs on its own.

Bar chart of characters sent on the first model call: Pi stock 5,521, SpecPi default 12,287, OpenCode 31,043, DeepSeek Harness 31,743, SpecPi enabled 36,935, Codex CLI 41,616, Oh My Pi 66,708, Claude Code 90,460.

Measured tool schemas + system/developer instructions · node scripts/measure-context.mjs --chart --omp=<path to Oh My Pi's cli.js> --oc=<path to OpenCode's binary> --dsh=<path to the DeepSeek Harness bin> · Recorded measurements and package pins · Method and caveats

The gap between the two SpecPi bars comes from a few separate switches, so the enabled tools are also measured group by group. For example, the fourteen browser QA tools add up to less than the four web access tools:

Bar chart of tool-schema characters each capability adds: Pi built-ins 2,896, Improvement loop 1,601, Capability request 995, Background jobs 728, Goals 1,315, Browser QA 8,234, Delegation 4,453, Web access 11,298. Browser QA, Delegation, Web access are hidden until switched on.

Every tool in the measured request belongs to exactly one group · Leaving all three opt-in groups hidden keeps 23,797 characters of tool schema out of every request

Harness evaluations

The chart above counts characters. It says nothing about what a harness costs to actually use, or whether it finishes the job. That is what the evals are for: Terminal-Bench 2.0, the same model and the same frozen price list, with only the harness changing.

947 scored attempts on two benchmarks and 6 harnesses, all on deepseek-v4.1-flash. SpecPi is the published 0.33.0 release, with the experimental Jev layer off.

SWE-bench Verified, 12 tasks × 3: SpecPi solved 34/36 against Pi's 34/36 (p = 1.00), with 2% more prompt tokens and 1% less cost per attempt.

Terminal-Bench 2.0, 731 attempts across 20 tasks:

Harness Solved Rate Cost/attempt Prompt tokens Cache hit
OpenCode 29/39 0.744 $0.0124 494,173 96.5%
SpecPi 62/77 0.805 $0.0144 407,416 92.4%
Pi (base) 164/230 0.713 $0.0155 436,214 92.9%
DeepSeek Harness 21/38 0.553 $0.0226 1,126,586 95.4%
Oh My Pi 115/151 0.762 $0.0237 1,108,076 96.5%
Claude Code 81/112 0.723 $0.0273 670,828 95.4%

SpecPi and Pi ran side by side in the 24 Sep · a sitting. SpecPi solved 30/38 against Pi's 25/39 (Fisher p = 0.21), sending 16% fewer prompt tokens and costing 22% less per attempt. On sanitize-git-repo, with that sitting's extra attempts, SpecPi solved 10/10 against 3/10 (p = 0.003); fix-git was 10/10 for both.

Overall solve rate is a different matter: one sitting cannot rank harnesses here. Bare Pi, on unchanged software and the same thirteen tasks, spans 56-77% across 5 sittings, a wider gap than any measured between two harnesses. Pooled across sittings, SpecPi leads DeepSeek Harness (p = 0.007) and Oh My Pi leads DeepSeek Harness (p = 0.015), but pooling sets one harness's sittings against another's. Cost is recomputed from recorded tokens against a dated price file, never taken from a harness's self-report.

This run is still in progress. See the results and brief method on the evaluations page. The table above is regenerated from the run data by node scripts/tb2-site.mjs, so it cannot drift from the published figures.

Install

Requires Node.js 22.19+, Git, npm, and an existing Pi installation on PATH. The base is tested with Pi 1.0.0.

npm install --global specpi@latest
specpi plan
specpi install
specpi doctor

plan shows what will change without modifying anything. Restart Pi after install.

Install also turns on Pi's built-in codemode tool, which lets the model run a short sandboxed script that calls several tools at once. It does this by adding +codemode to defaultTools in Pi's settings.json, only on Pi 0.99 or later. To keep it off, put -codemode there instead; SpecPi leaves an existing choice alone, and uninstall removes only the entry it added.

Full setup options, package details, and requirements: website.

Where things live

Packages The eight pinned packages and what each provides
Scope control /scope commands and drift monitoring
Command guard LANCET, the local classifier that checks shell commands before they run
Improvement loop Local wishlist, /harness-improvement, and retirement with evidence
SpecPi Chat VS Code frontend and VSIX install · Chat guide
Updating Update, uninstall, and migration notes

Development

npm install --ignore-scripts --omit=peer --no-package-lock
node --test tests/workflow-controls.test.mjs tests/workflow-controls-extension.test.mjs
npm run check

Installer tests use disposable Pi directories — never test against a live Pi installation. Publication follows the release procedure.

Security model · Third-party components · Release notes · MIT License

About

A minimal, explicit, provider-safe harness for the Pi coding agent

Topics

Resources

Security policy

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages