Skip to content

Latest commit

 

History

38 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AgentProbe

A ReAct agent written from scratch, with a deterministic failure taxonomy, a multi-model benchmark and a decision-graph view of every run.

CI Python 3.10+ TypeScript strict Ruff License: MIT Last commit

Decision graph of the seeded failure run: a hallucinated tool, recovery, then a drifted final answer

Walkthrough (screenshots below) · How it works · Failure taxonomy · Run it locally · Why this exists

There is no live demo right now. The screenshots were captured in May 2026 from a Railway + Vercel deployment. The Railway backend has since been removed, so the old Vercel front-end loads an empty shell and is not linked here. Everything below runs locally with docker compose up. docs/DEPLOY.md lists what a redeploy needs.

The front-end also has a keyless replay mode (see Replay mode): built with NEXT_PUBLIC_DATA_MODE=replay, it needs no backend and no API key and shows data committed under frontend/public/replay/. It shows one recorded benchmark run (anthropic/claude-haiku-4-5-20251001, 2026-09-24; see Results) with every case's score and trace, plus the 52 benchmark cases and the two hand-written demo runs. It has not been deployed.

Why this exists

Agent demos usually work; agent deployments often don't. The failure is rarely one bad model output. It is the gap between what the model decided and what the loop executed: a tool that does not exist, an action the parser could not read, the same call repeated until the step budget runs out.

AgentProbe makes that gap visible. The ReAct loop is written by hand (orchestrator.py and parser.py, about 420 lines), so every step of the reasoning path can be read and tested, with no framework in between. Each failure is tagged by a deterministic rule rather than an LLM judge, and each run is stored so it can be replayed, compared across models and aggregated. The longer argument is in docs/WHY.md.

What it solves

Layer Problem How AgentProbe answers it
Loop Framework agents hide the prompt, the parse and the dispatch. A from-scratch ReAct loop (Thought → Action → Observation) with a regex parser and an explicit tool registry.
Detection "The agent went wrong" is not a diagnosis. Deterministic checks tag each step with a failure type: tool lookup, regex parse, [ERROR] prefix, step and size limits.
Persistence A failure seen once in a terminal cannot be studied. Every run, step and failure goes to SQL (SQLite in dev, Postgres in prod) and can be replayed as the original SSE stream.
Benchmark One model's good run says nothing about another model. 52 cases in 5 categories, run per model and scored on answer, tools, efficiency and reliability.
View Long traces are hard to read as text. A decision graph per run, with failures drawn on the failing node, plus analytics over all stored runs.

Architecture

flowchart LR
    ui["Next.js 16 front-end<br/>playground, runs, benchmarks, analytics, compare, prompts"] -- "REST + SSE" --> api["FastAPI routes<br/>11 route modules, 3 middleware"]
    api --> orch["AgentOrchestrator<br/>ReAct loop"]
    orch --> parser["Parser<br/>Thought / Action / Final Answer"]
    orch --> reg["Tool registry<br/>calculator, web_search, think,<br/>read_file, memory, custom"]
    orch --> prov["LLM providers<br/>Groq, Ollama, OpenAI,<br/>Anthropic, Google"]
    parser --> det["Failure checks"]
    reg --> det
    det --> db[("SQL: runs, steps, failures<br/>11 tables, Alembic")]
    orch --> db
    eval["EvalHarness + ScoringEngine<br/>52 benchmark cases"] --> orch
    eval --> db
    db --> an["AnalyticsService"]
    an --> api
Loading

The backend follows a domain / application / infrastructure split: entities and ports have no external dependencies, services hold the loop, scoring and analytics, and providers, persistence, tools and routes sit at the edge.

The ReAct loop

flowchart TD
    start(["User query"]) --> loop{"Step < max?"}
    loop -- "no" --> timeout["max_steps_exceeded"]
    loop -- "yes" --> ctx{"Context under<br/>char limit?"}
    ctx -- "no" --> overflow["context_overflow"]
    ctx -- "yes" --> llm["Call LLM"]
    llm -- "empty or API error" --> empty["empty_response"]
    llm --> parse["Parse output"]
    parse -- "Final Answer" --> done(["Persist run"])
    parse -- "unparseable" --> malformed["malformed_action"] --> loop
    parse -- "Action" --> exists{"Tool in registry?"}
    exists -- "no" --> halluc["hallucinated_tool"] --> dup
    exists -- "yes" --> dup{"Same tool + input<br/>in last 4 steps?"}
    dup -- "yes" --> repeated["repeated_action"] --> loop
    dup -- "no" --> exec["Execute tool<br/>(unknown tool returns [ERROR])"] --> obs{"Observation<br/>has [ERROR]?"}
    obs -- "yes" --> toolerr["tool_execution_error"] --> loop
    obs -- "no" --> loop
Loading

Failure taxonomy

Eight failure types are defined in domain/entities/step.py. Seven are detected at runtime in orchestrator.py; goal_drift is not yet.

Failure What happened How it is detected
hallucinated_tool The model named a tool that does not exist Tool name not in the registry
malformed_action Output could not be parsed Regex parse of Action: / Action Input: fails
tool_execution_error The tool failed [ERROR] in the observation
max_steps_exceeded No final answer within the budget Step counter reaches the limit
context_overflow The prompt grew too large Character count over a configured limit
repeated_action The agent is looping Same tool and input within the last 4 steps
empty_response The model returned nothing Empty output, or the provider call raised
goal_drift The answer does not address the query No runtime detector. It appears only in the hand-written demo run

Walkthrough

These are screenshots, not a running app. They were captured with scripts/capture-screenshots.mjs against the May 2026 deployment.

The decision-graph and side-panel images show demo-fail-001, a hand-written demo run inserted by backend/scripts/seed_demo_runs.py, not a live model run. Its goal_drift tag was set by that script. The analytics image shows the 8 runs stored in that deployment's database at capture time, made with two Llama models (llama-3.1-8b-instant, llama-3.3-70b-versatile). It is not a benchmark result.

Side panel showing the hallucinated_tool step

Hallucinated tool. The red WEATHER_FORECAST action node: the side panel shows the rule (tool not in registry), the arguments and the step index.

Side panel showing the goal_drift tag on the final answer

Goal drift (seeded). The final answer is about Lyon's weather, not its population. The tag comes from the seed script, not from a detector.

List view of the same run

List view. The same run in order, showing the raw [ERROR] observation text.

Analytics dashboard over 8 stored runs

Analytics. Failure distribution and per-model comparison, computed with SQL aggregates over the stored runs (8 at capture time).

Features

Feature What it does
Decision graph Each stored run as a graph of reasoning steps, with failures on the failing node; toggle to a list
Playground Type a query and watch the agent reason step by step over SSE
Benchmark 52 cases (backend/data/benchmark_cases.json): math 12, search 10, reasoning 10, multi-tool 10, edge cases 10
Compare Two models on the same query, side by side
5 providers Groq, Ollama, OpenAI, Anthropic, Google, each marked available only when its key is set (Ollama always)
Replay Replay any stored run as an SSE stream
Static replay mode A front-end build that reads committed JSON instead of the API: runs, cases and recorded suites only
Analytics Failure distribution, per-model stats and a failure heatmap
Custom tools and prompts HTTP or static tools and system-prompt templates, created from the UI
Agent memory Key-value memory across runs through save and recall tools
Auth JWT and API keys behind the AUTH_ENABLED flag; custom tools, prompts and memory are scoped per user (runs are not)
Export CSV for runs and benchmarks, PDF for benchmarks

Benchmark scoring

application/services/scoring.py scores each case as a weighted sum: answer 40% (exact, substring, then keyword overlap with the expected answer), tools 20% (Jaccard similarity with the expected tools), efficiency 20% (full marks up to 3 steps, then decaying) and reliability 20% (1 if the run recorded no failures, 0 otherwise). A case passes when the answer score and the composite are both at least 0.5.

Results

One recorded run is committed: frontend/public/replay/recordings/2026-09-24T1621Z-full-anthropic-claude-haiku-4-5-20251001.json, recorded on 2026-09-24 from commit ae5bdcd. It is a single run with no seed variance, so the numbers below are one sample, not a mean over repeats. web_search used its mock result (no TAVILY_API_KEY), so the search scores say little about real search. No case hit a provider error.

Model: anthropic/claude-haiku-4-5-20251001, all 52 cases, max_steps=10.

Category Passed Mean score
Math 10 / 12 0.773
Search 10 / 10 0.796
Reasoning 10 / 10 0.737
Multi-tool 3 / 10 0.480
Edge cases 9 / 10 0.763
All 42 / 52 0.712

Answer judged correct on 42 cases and tools on 37; the run averaged 6.6 steps per case. The failure taxonomy tagged one failure in the whole run: tool_execution_error × 1 (edge-002, "Divide 10 by 0", where the calculator returned its division-by-zero error); every other type was 0. edge-003 (empty query) passed with 0.96: the agent now asks for a question without calling the model.

Of the 10 failed cases, 6 have an answer that looks right but that the substring and keyword scorer rejects on formatting: thousands separators (12,600 vs 12600 in math-009, also multi-005, multi-007, multi-008), rounding (8.05 / 8.0467 vs 8.047 in math-004) and a range (1,127 vs 1000+ in multi-009). The other 4 (multi-004, multi-006, multi-010, edge-005) differ from the expected answer. The scores above are what the scorer gave; they were not adjusted.

Usage, as reported by the provider: 96 LLM calls, 62,741 input and 14,839 output tokens. At the list price in backend/data/model_prices.json ($1 / $5 per 1M tokens) that is $0.1369, a list-price estimate computed from the reported tokens, not a billed amount. The pre-run cost estimate is in docs/BENCHMARK-COST.md.

Other verified checks:

Check Result Source
Backend tests 215 passed pytest in backend/
Backend lint and format clean ruff check, ruff format --check
Frontend unit tests 26 passed npm test (Vitest, replay loader and committed bundle)
Frontend type check clean npm run typecheck
Frontend lint 0 errors (4 unused-variable warnings) npm run lint
Frontend production build passes, live and replay mode npm run build, and with NEXT_PUBLIC_DATA_MODE=replay

Getting started

Docker

git clone https://github.com/soneeee22000/AgentProbe.git
cd AgentProbe

cat > .env << EOF
GROQ_API_KEY=your_key_here
TAVILY_API_KEY=your_key_here
EOF

docker compose up --build
# Frontend http://localhost:3000 · Backend http://localhost:8000 · API docs http://localhost:8000/docs

Seed the two demo runs (no API key needed), then open http://localhost:3000/runs/demo-fail-001 and /runs/demo-happy-001:

docker compose exec backend python -m scripts.seed_demo_runs

Without Docker

cd backend
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
cp .env.example .env
mkdir workspace
python -m scripts.seed_demo_runs
uvicorn main:app --reload --port 8000

cd ../frontend
npm install
npm run dev

A Groq or Google AI Studio key is enough for the playground and benchmarks; Ollama runs with no key. OpenAI and Anthropic keys are optional. Without a Tavily key, web_search returns a mock result.

Replay mode (no backend, no key)

cd frontend
NEXT_PUBLIC_DATA_MODE=replay npm run build
npx next start

The build reads public/replay/manifest.json and the files it lists instead of calling the API. Every page shows a banner saying what the data is:

  • With no recording committed: "Static replay with no recorded benchmark run yet. The runs shown are hand-written demo data, not model output."
  • After a recording is committed (the state today): "Replay of a recorded run on <date>, models <provider/model ids>".

Runs, run detail (graph, list and step-by-step replay), benchmark cases, recorded suites and analytics work. The playground, compare, prompts, suite runner and exports need a backend, so they show a notice instead. Analytics use recorded runs when any exist; otherwise they use the two hand-written runs, with a note that these are not benchmark results.

Two backend scripts produce the bundle:

cd backend
python -m scripts.export_replay_bundle                  # seed-runs.json, cases.json, manifest.json
python -m scripts.record_benchmark --estimate           # cost estimate, no calls
python -m scripts.record_benchmark --smoke --dry-run    # plan for 3 cases on groq/llama-3.1-8b-instant, no calls
python -m scripts.record_benchmark --smoke              # the smoke run (needs GROQ_API_KEY)
python -m scripts.record_benchmark --target groq/llama-3.1-8b-instant --target google/gemini-2.5-flash-lite

A recording stores each suite, each case's score and the full step trace of every run, plus the date, the model ids, the git commit, a hash of the case file and whether web_search used Tavily or its mock. A failed provider call (rate limit, bad key, unknown model id) is scored as a failed case, so the script does not publish a run in which any case hit one: it saves it to backend/recordings-local/ and exits with code 2. Each run also stores the input and output tokens the provider reported, per case and per suite, with a cost computed from those tokens and the list price. --max-spend-usd X refuses a run whose ceiling estimate passes X and stops making calls once the computed spend reaches it. Read docs/BENCHMARK-COST.md before a paid run.

Checks

cd backend && ruff check src/ tests/ && ruff format --check src/ tests/ && pytest --cov=src/agentprobe
cd frontend && npm run lint && npm run typecheck && npm test && npm run build

The API reference and database schema are in docs/API.md.

Project structure

AgentProbe/
├── backend/
│   ├── src/agentprobe/
│   │   ├── domain/            # entities, FailureType, port interfaces
│   │   ├── application/       # orchestrator, parser, eval harness, scoring, analytics, auth, export
│   │   └── infrastructure/    # FastAPI routes + middleware, 5 providers, SQLAlchemy (11 tables), tools
│   ├── scripts/               # seed_demo_runs, maybe_seed, render_demo_graph, record_benchmark, export_replay_bundle
│   ├── data/                  # benchmark_cases.json (52 cases), model_prices.json (estimate only)
│   ├── alembic/               # 3 migrations
│   ├── tests/                 # unit, integration, api (176 tests)
│   ├── Dockerfile, railway.toml
│   └── pyproject.toml
├── frontend/
│   ├── src/app/               # 8 pages: playground, runs, run detail, benchmarks, suite detail, analytics, compare, prompts
│   ├── src/components/        # UI components
│   ├── src/lib/               # api.ts (API client, SSE reader), replay.ts (replay loader, Vitest)
│   ├── public/replay/         # committed replay bundle: manifest, cases, hand-written runs, recordings
│   └── e2e/                   # Playwright specs (4)
├── scripts/capture-screenshots.mjs
├── docs/                      # WHY, API, DEPLOY, screenshots
└── docker-compose.yml

Limitations

  • No live deployment. The Railway backend is gone (Application not found), and the Vercel front-end has no data without it. Screenshots are from May 2026.
  • The showcase trace is hand-written. demo-fail-001 and demo-happy-001 are seed data, written to show the graph view, not recorded from a model.
  • goal_drift has no detector. It is in the taxonomy and the seed data only.
  • Detection is shallow by design. Rules are string and lookup checks. They catch structural failures, not wrong reasoning that is well formed. A provider exception is recorded as empty_response.
  • One benchmark run, one model. The committed results are a single Haiku 4.5 run with mock search and no repeats, so they carry no variance estimate and say nothing about other models.
  • Replay mode only replays. It cannot run the agent. It shows exactly what was committed: the recorded Haiku run, the hand-written demo runs and the case list.
  • The cost estimate is approximate. Tokens are estimated from characters (4 per token) with assumed turn and observation sizes, and the two Groq prices could not be confirmed on Groq's site on 2026-09-24.
  • Answer scoring is keyword overlap, so correct paraphrases can score low and wrong answers that reuse the expected words can score high.
  • mypy and E2E are not gating. CI runs both with continue-on-error; mypy currently reports 69 errors, and the Playwright E2E job is not required to pass.
  • Postgres behaviour was tested only in that deployment. CI tests run on in-memory SQLite.

Roadmap

  • Implement a runtime goal_drift check, tests first, and stop relying on the seed script for it.
  • Record the 52-case benchmark on more providers, with real Tavily search and several repeats per model.
  • Make the E2E specs use a mock provider, and add an E2E check of the replay build.
  • Clear the mypy errors and make mypy and E2E blocking in CI.
  • Redeploy, if wanted, following docs/DEPLOY.md.

License

MIT. See LICENSE.

Author

Pyae Sone (Seon) · github.com/soneeee22000

About

A ReAct agent written from scratch, with a deterministic failure taxonomy and a 52-case evaluation harness. Recorded benchmark: Claude Haiku 4.5, 42/52. FastAPI · SSE · TypeScript

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages