Skip to content

Repository files navigation

IntentGate

Intent-Consistency Verification at the Action Gate: A Request-Derived Alternative to Formal Contract Gating for Tool-Using LLM Agents

LLM Security · Agentic AI · AI Safety — Thesis Project (2026) · github.com/Atik203/IntentGate

MCP-style agents now execute real actions (email, payments, code, files). Benchmarks show systemic hijack vulnerability — AgentDojo 629 cases, InjecAgent 1,054 cases (GPT-4 24%→47% ASR), MCPTox 1,312 cases on 45 live servers (up to 72.8% ASR, <3% refusal), ASB 84.3% mixed ASR — yet defenses are siloed or require hand-authored contracts (ToolGate, arXiv 2601.04688v1, never tested adversarially).

This project builds a model-agnostic middleware gate that derives an intent contract automatically from the user's original request and blocks/escalates any tool call inconsistent with that intent — evaluated head-to-head against an unprotected agent and a ToolGate reimplementation on the benchmarks ToolGate never tested.

Setup

Install and verify in ~5 minutes: see docs/setup.md (venv + pip install -e ".[dev]" + .env + pytest -q + smoke-test scripts).


Repository Structure

IntentGate\
├── README.md                      # this file
├── CONTRIBUTING.md                # branch model, PR rules, definition of done
├── SECURITY.md                    # vulnerability reporting + scope
├── CODE_OF_CONDUCT.md             # Contributor Covenant v2.1
├── CHANGELOG.md                   # Keep a Changelog; updated per PR
├── LICENSE                        # MIT
├── AGENTS.md                      # instructions for AI coding agents (+ LLM cache guidance)
├── roadmap.md                     # phased tracker (Phases 0–8, Gates 0–3, success criteria)
├── docs\
│   ├── README.md                  # docs hub (where everything lives + status snapshot)
│   ├── setup.md                   # install + verify instructions (venv, pip, .env, tests)
│   ├── blueprint.md               # single source of truth (design, Sec 0–18, indexed)
│   ├── references.bib             # BibTeX citation basis (10 verified papers)
│   └── experiments\               # experiment record: README hub + 01–04 reports (E1–E10)
├── pyproject.toml                 # intent-gate package (pip install -e .)
├── .github\                       # PR/issue templates + CI (pytest on dev/main)
├── configs\                       # intent_schema.json, parser_fewshots.json, thresholds.yaml, models.yaml
├── src\intent_gate\               # parser / scoring / gate / agent / baselines.toolgate / eval
├── harness\                       # run_injecagent.py, run_mcptox.py, compare_b2.py, common.py
├── scripts\                       # pilot_score_dist.py, check_parser.py, clone_benchmarks.ps1
├── tests\                         # pytest suite (parser, veto, no-bypass, ordering, metrics)
├── data\                          # (gitignored) benchmark clones + snapshots — scripts/clone_benchmarks.ps1
├── results\                       # (gitignored) JSONL traces + reports
├── literature_review\
│   ├── index.md                   # Master Comparison Matrix + Gap Map + Verification Log (5 anchors + Closest)
│   └── papers\
│       ├── 01-agentdojo-debenedetti-2024.md   # Anchor — dynamic 97-task/629-case injection harness
│       ├── 02-injecagent-zhan-2024.md         # Anchor — 1,054 indirect-injection cases
│       ├── 03-mcptox-wang-2025.md             # Anchor — 1,312 live-MCP poisoning cases
│       ├── 04-asb-zhang-2025.md               # Anchor — 10-scenario/27-method superset
│       └── 05-toolgate-liu-2026.md            # Closest — Hoare-contract gate (B2 baseline)
├── literature_review.md           # (local, gitignored) original 10-paper Q1–Q9 synthesis
├── mds\                           # (optional, gitignored) pasted paper markdowns for offline access
└── pdfs\                          # (optional, gitignored) paper PDFs

Branch Model (5-member team)

feature branch (feat/, fix/, docs/, exp/, <member>/) ──PR──> dev ──PR (tested+reviewed)──> main
  • main — production. Only merges from dev after all tests pass. Never commit directly.
  • dev — integration branch. All feature PRs target dev.
  • Feature branches start from dev, named after the feature or the member (e.g. feat/gate-veto, atik/toolgate-contracts).

Full rules: CONTRIBUTING.md. Security reports: SECURITY.md.


Literature Review

  • Index: literature_review/index.md — Master Matrix (5 verified 2026-08-24), Legend, Quick Triage, Gap Map, Verification Log
  • Papers: 5 detailed reviews in literature_review/papers/ — each follows the same template (badges, Summary, Relevant to Our Idea, Gap, Q1–Q9, Citation, Method, Results, Limitations, Comparison, Positioning, Reproducibility, Cross-References, Relevance to Thesis)
  • Blueprint: docs/blueprint.md — Phase 1–5, Sections 0–18 (indexed with anchors): design decisions, pipeline, data flow, models/tools, datasets, evaluation, edge cases, risks, roadmap, implementation order, supervisor/team explanations, expected outcomes, future work, Reviewer #2 critique
  • Roadmap: roadmap.md — phased tracker (Phases 0–8, Gates 0–3, success criteria, current status)
  • Docs hub: docs/README.md · Experiments: docs/experiments/README.md — status matrix, experiment index E1–E10, datasets, pinned models, artifact map, reproduce commands, supervisor Q&A

All 5 papers verified via full arXiv html (not snippets). BibTeX: docs/references.bib.


Core Idea (10-point summary)

  1. Title: Intent-Consistency Verification at the Action Gate: A Request-Derived Alternative to Formal Contract Gating
  2. Domain: LLM Security · Agentic AI · AI Safety
  3. Motivation: Shift from chat to action via MCP; OWASP excessive agency Top 10; formal gating exists but is manual and non-adversarial
  4. Problem: Injection, tool poisoning, and drift all converge to one observable — an agent executes an action the user never intended
  5. Gap: Single-vector filters vs. manual Hoare contracts (ToolGate, never on InjecAgent/MCPTox); no auto-derived, cross-vector gate
  6. Solution: Parse user request → structured intent contract (goals, tool categories, data scopes, side-effect limits) → middleware gate scores every proposed tool_call(name,params) via S = α·cosine(embed(contract), embed(call)) + (1-α)·rule_compliance with veto; decision allow / block / escalate via threshold τ (sweep)
  7. Novelty: Zero per-tool authoring, vector-agnostic (injection + poisoning), first adversarial test of contract-style gating
  8. Deliverables: Pip-installable wrapper, ToolGate reimplementation (Appendix G), evaluation report (ASR/FPR/latency/setup-cost), workshop/Findings paper draft
  9. Evaluation: B1 unprotected ReAct vs B2 ToolGate vs Ours on InjecAgent (1,054) + MCPTox (1,312, 45 servers) — metrics ASR, FPR, escalation rate, latency p95, setup cost (0 vs manual count); stretch multi-turn drift pilot (20 AgentDojo cases)
  10. Stack: Open LLM (Qwen/Llama via API), ReAct/MCP client, Python middleware, all-MiniLM-L6-v2 (CPU), no fine-tuning

Evaluation Plan (abridged from docs/blueprint.md:Section 9)

  • B1: Unprotected ReAct (no gate) — raw vulnerability floor
  • B2: ToolGate reimplementation (Hoare {P}T{Q} per Appendix G, symbolic world-state) — validated on ToolBench/MCP-Universe subset before adversarial runs
  • Metrics: ASR-valid, FPR, escalation rate, latency p50/p95, setup cost, utility retention on benign cases
  • Ablations: A1 semantic-only, A2 rule-only, A3 raw-request vs structured contract, A4 threshold/escalate sweep
  • Stats: Bootstrap 95% CI on ASR, paired McNemar per case, per-risk-category breakdown (MCPTox 11 categories, InjecAgent 17 tools)

Quick Start (planned)

# 1. Clone benchmarks (Week 1)
git clone https://github.com/uiuc-kang-lab/InjecAgent
git clone https://anonymous.4open.science/r/AAAI26-7C02  # MCPTox snapshot
git clone https://github.com/OceannTwT/ToolGate          # B2 reference
git clone https://github.com/ethz-spylab/agentdojo       # stretch

# 2. Install gate (after Weeks 5–8)
pip install -e .  # pip wrapper around agent tool executor

# 3. Run evaluation (Weeks 9–12)
python harness/run_injecagent.py --cases data/raw/InjecAgent/data/test_cases_dh_base.json --gate ours --threshold 0.6
python harness/run_mcptox.py --gate ours --snapshot v1 --report results/mcptox.json
python harness/compare_b2.py --cases data/raw/InjecAgent/data/test_cases_dh_base.json --report results/comparison.json

See docs/blueprint.md:Section 12 for 16-week roadmap (Weeks 1–2 pilot → Weeks 3–4 B2 → Weeks 5–8 gate → Weeks 9–12 full evaluation → Weeks 13–16 writing) and docs/blueprint.md:Section 13 for exact build order.


Publication Target

Realistic: security/agentic-AI workshop or Findings track (EMNLP/ACL Findings, USENIX-adjacent) — Q2 journal equivalent (e.g., Journal of Information Security and Applications, Applied Intelligence) is feasible with rigorous ASR/FPR/latency/setup-cost reporting and error taxonomy (see docs/blueprint.md:Section 18 Reviewer #2 checklist).


Verified Papers

# Paper Venue Role
1 AgentDojo (Debenedetti et al., 2024) NeurIPS D&B Anchor benchmark
2 InjecAgent (Zhan et al., 2024) Findings of ACL 2024 Anchor benchmark
3 MCPTox (Wang et al., 2025) arXiv → AAAI 2026 Anchor benchmark (live)
4 ASB (Zhang et al., 2025) ICLR 2025 Anchor superset
5 ToolGate (Liu et al., 2026) arXiv 2601.04688v1 Closest system (B2)

Full reviews: literature_review/papers/


Contributing

5-member team workflow: feature branch → dev → main. Read CONTRIBUTING.md before your first PR; CI runs pytest -q on dev and main. All contributors follow CODE_OF_CONDUCT.md. Report vulnerabilities privately per SECURITY.md (never in public issues).

License

MIT — Copyright (c) 2026 IntentGate contributors.


Citation

@misc{intentgate2026,
  title={Intent-Consistency Verification at the Action Gate: A Request-Derived Alternative to Formal Contract Gating for Tool-Using LLM Agents},
  year={2026},
  note={Thesis project — IntentGate, https://github.com/Atik203/IntentGate}
}

Last updated: 2026-09-16 — blueprint is the single source of truth, roadmap.md tracks status.

About

Intent-Consistency Verification at the Action Gate: A Request-Derived Alternative to Formal Contract Gating for Tool-Using LLM Agents

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages