Your AI agent made the tests green. Did it fix the code, or the tests?
A quality advisor for coding agents: git hooks and Stop hooks that raise a loud red flag on test tampering, secrets and skipped hooks, and hold one complexity bar on every changed function. They never block: the agent sees the evidence and decides. 12 languages, one config file, no server.
Nothing here stops the agent. No hook refuses a commit, a push, a shell command or the end of a turn. The plugin trusts the model the way a good reviewer trusts a colleague: it measures, shows its work and says plainly when something is dangerous. The agent decides, and says why.
Dangerous moves are red flags, printed first in every output:
RED FLAG: tamper/test-deleted tests/queue.test.ts:1 test file deleted -- a green run stops proving the code works; check: is the test change intended? restore it, or say why in the commit message
Red flags: a secret in the change or in the pushed history, a deleted, skipped or weakened test, a baseline refreshed with the code, red tests on push, and a git command that skips the hooks (--no-verify, commit -n, core.hooksPath). Everything else (complexity, CRAP, diff coverage, cycles, dead code, clones, glossary, AI attribution, Jev) is a finding or a note in the report.
- Advice, never a block (1.3.3). No hook stops a commit, a push, a shell command or the end of a turn. Dangerous moves are
RED FLAGlines, first in every output; the Stop report reaches the agent once, with its next message. - Accept a change by its own tests.
check --since <base> --testsruns every test that reaches the change through the project's imports, with coverage, and judges only the changed functions. - Prove a test catches a bug.
mutantmakes one exact mutation, runs the tests and restores the file byte for byte, even after Ctrl-C. - The report is about the pushed commit. When a push has touched tests to run, pre-push runs them only on the checked-out commit and says so otherwise.
- Calm on a shared machine. Test runs wait for a busy machine, Ctrl-C kills the whole test group, and a copy run by hand no longer takes the machine's git hooks.
Full notes: releases.
| The agent | The report |
|---|---|
Marked the failing test it.skip |
tamper/test-skipped tests/queue.test.ts:12 |
| Dropped two assertions | tamper/assertion-weakened tests/queue.test.ts:48 |
Swapped toEqual for toBeTruthy |
tamper/assertion-weakened tests/queue.test.ts:52 |
| Stubbed the module under test | note: tamper/mock-added tests/queue.test.ts:+64 |
| Refreshed the baseline with the code | tamper/baseline-touched baseline.json:1 |
| Shipped 50 lines, 31 covered | cov/diff: 62% (31/50 lines, minimum 80%) |
Deterministic: file and line in the report, red flags first. No model is asked for an opinion, and the commit still goes through; the agent answers the flag.
flowchart LR
A[git hooks] --> Q
B[Stop hook: Claude, Codex, pi, omp; Grok after global hook install] --> Q
C[OpenCode idle notice] --> Q
Q[one script] --> R[RED FLAG lines first: secrets, tampered tests, red tests on push, skipped hooks]
R --> D{new finding on a changed line?}
D -->|yes| F[FINDINGS, advice to the agent]
D -->|no| P[CLEAN]
F --> G[the commit, push or turn goes ahead; the agent decides]
P --> G
- One bar. Cyclomatic complexity 10, CRAP 30, 80% coverage of added lines, no new cycles, no new dead code, no secrets.
- Old debt is not news. A baseline snapshot keeps existing findings as debt. Only new findings on changed lines are reported.
- Seconds, not minutes.
checkruns in 3-19 s on a real repo. The full run with tests and coverage is a manual one: a debt measurement, not a verdict on a change. - Acceptance on the change.
check --since <base> --testsruns the change's touched tests with coverage and reports on the change with it.checkexits 1 on findings so a script or CI can decide; no hook turns that into a stop.mutantproves a test catches one exact bug and always restores the file. - Any language. Complexity from lizard, coverage from lcov, adapters for TypeScript, Python, Go, Rust, Java, Kotlin, C#, Swift, PHP, Ruby, C/C++, Dart.
Bun is required for the gate and the shell hooks. Install for your agent:
| Agent | Install |
|---|---|
| Claude Code | claude plugin marketplace add smixs/code-quality then claude plugin install code-quality@code-quality |
| Codex | codex plugin marketplace add smixs/code-quality then codex plugin add code-quality@code-quality; review and trust its hooks with /hooks |
| Grok 1.0.40 | grok plugin install smixs/code-quality --trust, then from this repository or its installed plugin root run bun scripts/quality.ts install-grok-hooks |
| pi | pi install git:github.com/smixs/code-quality |
| omp | omp plugin install github:smixs/code-quality |
| OpenCode V1 | bun add github:smixs/code-quality in a project with package.json, then add "plugin": ["file:./node_modules/code-quality"] and "skills": ["./node_modules/code-quality/skills"] to opencode.json |
| OpenCode V2 | Add "plugins": ["github:smixs/code-quality"] to opencode.json |
| Skill only | npx skills add smixs/code-quality (does not install hooks) |
Claude and Codex load Stop, UserPromptSubmit and PreToolUse hooks from the plugin. The Stop hook never holds the turn: the user sees the red flags at once, and the agent gets the full report once, with its next message (UserPromptSubmit). Grok 1.0.40 did not dispatch plugin hooks in a live check (total_hooks=0); install-grok-hooks writes ~/.grok/hooks/code-quality.json with both commands pointing to the current plugin root; Grok drops UserPromptSubmit context, so there the Stop report reaches the user only. Run it again after moving or updating that root; bun scripts/quality.ts uninstall-grok-hooks removes that file. $GROK_HOME overrides the Grok directory. pi and omp load the package extensions: a red flag is added to the shell command's result, the Stop report goes to the next turn. OpenCode adds the red flag after the shell command and writes the Stop report into the session on idle without a reply. OpenCode V1.18.32 reads the plugin's config.skills value but does not discover the skill from it. Add an explicit "skills": ["./node_modules/code-quality/skills"] entry to opencode.json when using a project Bun installation. In Git repositories with .quality.toml, git commit --no-verify, git commit -n, git push --no-verify, git -c core.hooksPath=... and git config core.hooksPath run as asked and get a red flag next to their output. Set [hooks] flag_bypass = false in .quality.toml to silence it (block_bypass, the old name, is still read).
To enable the plugin for a team repository, commit these project settings:
For Codex, commit .codex/config.toml:
[plugins."code-quality@code-quality"]
enabled = trueWire each Git repository with a .quality.toml and a baseline. From the plugin checkout or installed package root:
bun scripts/quality.ts install-hooks <repo>
bun scripts/quality.ts --repo <repo> --update-baselineinstall-hooks sets core.hooksPath to ~/.local/share/code-quality/git-hooks/, or $CODE_QUALITY_HOME/git-hooks/. The hooks run the newest plugin copy an agent invoked on this machine (its Stop or shell guard hook), so an agent session still on an older copy cannot take them back, and a copy run by hand from a clone or a worktree never takes them. Existing repository hooks are chained. See configuration.
Five calibrated yes/no questions about added test hunks (a textual test, an untested error path, a
weakened assertion, a mock that hides the change, a tautological property). It also asks whether each
changed source hunk has a test in the same diff, whether the change implements the task spec
([review] spec or QG_SPEC) and, opt-in, UX and AI-agent question packs plus your own project
questions. Notes only, like everything else here. Two ways to connect, whichever key you have:
[review] # TypeSafe directly, key from console.typesafe.ai/keys
jev = true # TYPESAFE_API_KEY in the environment[review] # or through OpenRouter
jev = true # OPENROUTER_API_KEY in the environment
jev_provider = "openrouter"Details, the model pins and the curl for each: Jev reference.
- Every rule, which ones are red flags and which are findings or notes
- Languages and adapters
- Configure
.quality.tomland hooks - Jev: an optional classifier for test and source hunks
- Acceptance: the runbook for accepting a change: a spec with a failure table, one parallel pass on a frozen candidate, one batched repair, two rounds
- SKILL.md, the full reference the agent reads
CRAP metric by Alberto Savoia and Brian Cunningham. Tamper taxonomy from TRACE (arXiv 2601.20103). Built on lizard, jscpd, ast-grep, gitleaks, osv-scanner, dependency-cruiser, knip, ESLint, Radon and TypeSafe Jev.
MIT. Copyright 2026 Sergey Shima.
