A ReAct agent written from scratch, with a deterministic failure taxonomy, a multi-model benchmark and a decision-graph view of every run.
Walkthrough (screenshots below) · How it works · Failure taxonomy · Run it locally · Why this exists
There is no live demo right now. The screenshots were captured in May 2026 from a Railway + Vercel deployment. The Railway backend has since been removed, so the old Vercel front-end loads an empty shell and is not linked here. Everything below runs locally with docker compose up. docs/DEPLOY.md lists what a redeploy needs.
The front-end also has a keyless replay mode (see Replay mode): built with NEXT_PUBLIC_DATA_MODE=replay, it needs no backend and no API key and shows data committed under frontend/public/replay/. It shows one recorded benchmark run (anthropic/claude-haiku-4-5-20251001, 2026-09-24; see Results) with every case's score and trace, plus the 52 benchmark cases and the two hand-written demo runs. It has not been deployed.
Agent demos usually work; agent deployments often don't. The failure is rarely one bad model output. It is the gap between what the model decided and what the loop executed: a tool that does not exist, an action the parser could not read, the same call repeated until the step budget runs out.
AgentProbe makes that gap visible. The ReAct loop is written by hand (orchestrator.py and parser.py, about 420 lines), so every step of the reasoning path can be read and tested, with no framework in between. Each failure is tagged by a deterministic rule rather than an LLM judge, and each run is stored so it can be replayed, compared across models and aggregated. The longer argument is in docs/WHY.md.
| Layer | Problem | How AgentProbe answers it |
|---|---|---|
| Loop | Framework agents hide the prompt, the parse and the dispatch. | A from-scratch ReAct loop (Thought → Action → Observation) with a regex parser and an explicit tool registry. |
| Detection | "The agent went wrong" is not a diagnosis. | Deterministic checks tag each step with a failure type: tool lookup, regex parse, [ERROR] prefix, step and size limits. |
| Persistence | A failure seen once in a terminal cannot be studied. | Every run, step and failure goes to SQL (SQLite in dev, Postgres in prod) and can be replayed as the original SSE stream. |
| Benchmark | One model's good run says nothing about another model. | 52 cases in 5 categories, run per model and scored on answer, tools, efficiency and reliability. |
| View | Long traces are hard to read as text. | A decision graph per run, with failures drawn on the failing node, plus analytics over all stored runs. |
flowchart LR
ui["Next.js 16 front-end<br/>playground, runs, benchmarks, analytics, compare, prompts"] -- "REST + SSE" --> api["FastAPI routes<br/>11 route modules, 3 middleware"]
api --> orch["AgentOrchestrator<br/>ReAct loop"]
orch --> parser["Parser<br/>Thought / Action / Final Answer"]
orch --> reg["Tool registry<br/>calculator, web_search, think,<br/>read_file, memory, custom"]
orch --> prov["LLM providers<br/>Groq, Ollama, OpenAI,<br/>Anthropic, Google"]
parser --> det["Failure checks"]
reg --> det
det --> db[("SQL: runs, steps, failures<br/>11 tables, Alembic")]
orch --> db
eval["EvalHarness + ScoringEngine<br/>52 benchmark cases"] --> orch
eval --> db
db --> an["AnalyticsService"]
an --> api
The backend follows a domain / application / infrastructure split: entities and ports have no external dependencies, services hold the loop, scoring and analytics, and providers, persistence, tools and routes sit at the edge.
flowchart TD
start(["User query"]) --> loop{"Step < max?"}
loop -- "no" --> timeout["max_steps_exceeded"]
loop -- "yes" --> ctx{"Context under<br/>char limit?"}
ctx -- "no" --> overflow["context_overflow"]
ctx -- "yes" --> llm["Call LLM"]
llm -- "empty or API error" --> empty["empty_response"]
llm --> parse["Parse output"]
parse -- "Final Answer" --> done(["Persist run"])
parse -- "unparseable" --> malformed["malformed_action"] --> loop
parse -- "Action" --> exists{"Tool in registry?"}
exists -- "no" --> halluc["hallucinated_tool"] --> dup
exists -- "yes" --> dup{"Same tool + input<br/>in last 4 steps?"}
dup -- "yes" --> repeated["repeated_action"] --> loop
dup -- "no" --> exec["Execute tool<br/>(unknown tool returns [ERROR])"] --> obs{"Observation<br/>has [ERROR]?"}
obs -- "yes" --> toolerr["tool_execution_error"] --> loop
obs -- "no" --> loop
Eight failure types are defined in domain/entities/step.py. Seven are detected at runtime in orchestrator.py; goal_drift is not yet.
| Failure | What happened | How it is detected |
|---|---|---|
hallucinated_tool |
The model named a tool that does not exist | Tool name not in the registry |
malformed_action |
Output could not be parsed | Regex parse of Action: / Action Input: fails |
tool_execution_error |
The tool failed | [ERROR] in the observation |
max_steps_exceeded |
No final answer within the budget | Step counter reaches the limit |
context_overflow |
The prompt grew too large | Character count over a configured limit |
repeated_action |
The agent is looping | Same tool and input within the last 4 steps |
empty_response |
The model returned nothing | Empty output, or the provider call raised |
goal_drift |
The answer does not address the query | No runtime detector. It appears only in the hand-written demo run |
These are screenshots, not a running app. They were captured with scripts/capture-screenshots.mjs against the May 2026 deployment.
The decision-graph and side-panel images show demo-fail-001, a hand-written demo run inserted by backend/scripts/seed_demo_runs.py, not a live model run. Its goal_drift tag was set by that script. The analytics image shows the 8 runs stored in that deployment's database at capture time, made with two Llama models (llama-3.1-8b-instant, llama-3.3-70b-versatile). It is not a benchmark result.
| Feature | What it does |
|---|---|
| Decision graph | Each stored run as a graph of reasoning steps, with failures on the failing node; toggle to a list |
| Playground | Type a query and watch the agent reason step by step over SSE |
| Benchmark | 52 cases (backend/data/benchmark_cases.json): math 12, search 10, reasoning 10, multi-tool 10, edge cases 10 |
| Compare | Two models on the same query, side by side |
| 5 providers | Groq, Ollama, OpenAI, Anthropic, Google, each marked available only when its key is set (Ollama always) |
| Replay | Replay any stored run as an SSE stream |
| Static replay mode | A front-end build that reads committed JSON instead of the API: runs, cases and recorded suites only |
| Analytics | Failure distribution, per-model stats and a failure heatmap |
| Custom tools and prompts | HTTP or static tools and system-prompt templates, created from the UI |
| Agent memory | Key-value memory across runs through save and recall tools |
| Auth | JWT and API keys behind the AUTH_ENABLED flag; custom tools, prompts and memory are scoped per user (runs are not) |
| Export | CSV for runs and benchmarks, PDF for benchmarks |
application/services/scoring.py scores each case as a weighted sum: answer 40% (exact, substring, then keyword overlap with the expected answer), tools 20% (Jaccard similarity with the expected tools), efficiency 20% (full marks up to 3 steps, then decaying) and reliability 20% (1 if the run recorded no failures, 0 otherwise). A case passes when the answer score and the composite are both at least 0.5.
One recorded run is committed: frontend/public/replay/recordings/2026-09-24T1621Z-full-anthropic-claude-haiku-4-5-20251001.json, recorded on 2026-09-24 from commit ae5bdcd. It is a single run with no seed variance, so the numbers below are one sample, not a mean over repeats. web_search used its mock result (no TAVILY_API_KEY), so the search scores say little about real search. No case hit a provider error.
Model: anthropic/claude-haiku-4-5-20251001, all 52 cases, max_steps=10.
| Category | Passed | Mean score |
|---|---|---|
| Math | 10 / 12 | 0.773 |
| Search | 10 / 10 | 0.796 |
| Reasoning | 10 / 10 | 0.737 |
| Multi-tool | 3 / 10 | 0.480 |
| Edge cases | 9 / 10 | 0.763 |
| All | 42 / 52 | 0.712 |
Answer judged correct on 42 cases and tools on 37; the run averaged 6.6 steps per case. The failure taxonomy tagged one failure in the whole run: tool_execution_error × 1 (edge-002, "Divide 10 by 0", where the calculator returned its division-by-zero error); every other type was 0. edge-003 (empty query) passed with 0.96: the agent now asks for a question without calling the model.
Of the 10 failed cases, 6 have an answer that looks right but that the substring and keyword scorer rejects on formatting: thousands separators (12,600 vs 12600 in math-009, also multi-005, multi-007, multi-008), rounding (8.05 / 8.0467 vs 8.047 in math-004) and a range (1,127 vs 1000+ in multi-009). The other 4 (multi-004, multi-006, multi-010, edge-005) differ from the expected answer. The scores above are what the scorer gave; they were not adjusted.
Usage, as reported by the provider: 96 LLM calls, 62,741 input and 14,839 output tokens. At the list price in backend/data/model_prices.json ($1 / $5 per 1M tokens) that is $0.1369, a list-price estimate computed from the reported tokens, not a billed amount. The pre-run cost estimate is in docs/BENCHMARK-COST.md.
Other verified checks:
| Check | Result | Source |
|---|---|---|
| Backend tests | 215 passed | pytest in backend/ |
| Backend lint and format | clean | ruff check, ruff format --check |
| Frontend unit tests | 26 passed | npm test (Vitest, replay loader and committed bundle) |
| Frontend type check | clean | npm run typecheck |
| Frontend lint | 0 errors (4 unused-variable warnings) | npm run lint |
| Frontend production build | passes, live and replay mode | npm run build, and with NEXT_PUBLIC_DATA_MODE=replay |
git clone https://github.com/soneeee22000/AgentProbe.git
cd AgentProbe
cat > .env << EOF
GROQ_API_KEY=your_key_here
TAVILY_API_KEY=your_key_here
EOF
docker compose up --build
# Frontend http://localhost:3000 · Backend http://localhost:8000 · API docs http://localhost:8000/docsSeed the two demo runs (no API key needed), then open http://localhost:3000/runs/demo-fail-001 and /runs/demo-happy-001:
docker compose exec backend python -m scripts.seed_demo_runscd backend
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
cp .env.example .env
mkdir workspace
python -m scripts.seed_demo_runs
uvicorn main:app --reload --port 8000
cd ../frontend
npm install
npm run devA Groq or Google AI Studio key is enough for the playground and benchmarks; Ollama runs with no key. OpenAI and Anthropic keys are optional. Without a Tavily key, web_search returns a mock result.
cd frontend
NEXT_PUBLIC_DATA_MODE=replay npm run build
npx next startThe build reads public/replay/manifest.json and the files it lists instead of calling the API. Every page shows a banner saying what the data is:
- With no recording committed: "Static replay with no recorded benchmark run yet. The runs shown are hand-written demo data, not model output."
- After a recording is committed (the state today): "Replay of a recorded run on
<date>, models<provider/model ids>".
Runs, run detail (graph, list and step-by-step replay), benchmark cases, recorded suites and analytics work. The playground, compare, prompts, suite runner and exports need a backend, so they show a notice instead. Analytics use recorded runs when any exist; otherwise they use the two hand-written runs, with a note that these are not benchmark results.
Two backend scripts produce the bundle:
cd backend
python -m scripts.export_replay_bundle # seed-runs.json, cases.json, manifest.json
python -m scripts.record_benchmark --estimate # cost estimate, no calls
python -m scripts.record_benchmark --smoke --dry-run # plan for 3 cases on groq/llama-3.1-8b-instant, no calls
python -m scripts.record_benchmark --smoke # the smoke run (needs GROQ_API_KEY)
python -m scripts.record_benchmark --target groq/llama-3.1-8b-instant --target google/gemini-2.5-flash-liteA recording stores each suite, each case's score and the full step trace of every run, plus the date, the model ids, the git commit, a hash of the case file and whether web_search used Tavily or its mock. A failed provider call (rate limit, bad key, unknown model id) is scored as a failed case, so the script does not publish a run in which any case hit one: it saves it to backend/recordings-local/ and exits with code 2. Each run also stores the input and output tokens the provider reported, per case and per suite, with a cost computed from those tokens and the list price. --max-spend-usd X refuses a run whose ceiling estimate passes X and stops making calls once the computed spend reaches it. Read docs/BENCHMARK-COST.md before a paid run.
cd backend && ruff check src/ tests/ && ruff format --check src/ tests/ && pytest --cov=src/agentprobe
cd frontend && npm run lint && npm run typecheck && npm test && npm run buildThe API reference and database schema are in docs/API.md.
AgentProbe/
├── backend/
│ ├── src/agentprobe/
│ │ ├── domain/ # entities, FailureType, port interfaces
│ │ ├── application/ # orchestrator, parser, eval harness, scoring, analytics, auth, export
│ │ └── infrastructure/ # FastAPI routes + middleware, 5 providers, SQLAlchemy (11 tables), tools
│ ├── scripts/ # seed_demo_runs, maybe_seed, render_demo_graph, record_benchmark, export_replay_bundle
│ ├── data/ # benchmark_cases.json (52 cases), model_prices.json (estimate only)
│ ├── alembic/ # 3 migrations
│ ├── tests/ # unit, integration, api (176 tests)
│ ├── Dockerfile, railway.toml
│ └── pyproject.toml
├── frontend/
│ ├── src/app/ # 8 pages: playground, runs, run detail, benchmarks, suite detail, analytics, compare, prompts
│ ├── src/components/ # UI components
│ ├── src/lib/ # api.ts (API client, SSE reader), replay.ts (replay loader, Vitest)
│ ├── public/replay/ # committed replay bundle: manifest, cases, hand-written runs, recordings
│ └── e2e/ # Playwright specs (4)
├── scripts/capture-screenshots.mjs
├── docs/ # WHY, API, DEPLOY, screenshots
└── docker-compose.yml
- No live deployment. The Railway backend is gone (
Application not found), and the Vercel front-end has no data without it. Screenshots are from May 2026. - The showcase trace is hand-written.
demo-fail-001anddemo-happy-001are seed data, written to show the graph view, not recorded from a model. goal_drifthas no detector. It is in the taxonomy and the seed data only.- Detection is shallow by design. Rules are string and lookup checks. They catch structural failures, not wrong reasoning that is well formed. A provider exception is recorded as
empty_response. - One benchmark run, one model. The committed results are a single Haiku 4.5 run with mock search and no repeats, so they carry no variance estimate and say nothing about other models.
- Replay mode only replays. It cannot run the agent. It shows exactly what was committed: the recorded Haiku run, the hand-written demo runs and the case list.
- The cost estimate is approximate. Tokens are estimated from characters (4 per token) with assumed turn and observation sizes, and the two Groq prices could not be confirmed on Groq's site on 2026-09-24.
- Answer scoring is keyword overlap, so correct paraphrases can score low and wrong answers that reuse the expected words can score high.
- mypy and E2E are not gating. CI runs both with
continue-on-error; mypy currently reports 69 errors, and the Playwright E2E job is not required to pass. - Postgres behaviour was tested only in that deployment. CI tests run on in-memory SQLite.
- Implement a runtime
goal_driftcheck, tests first, and stop relying on the seed script for it. - Record the 52-case benchmark on more providers, with real Tavily search and several repeats per model.
- Make the E2E specs use a mock provider, and add an E2E check of the replay build.
- Clear the mypy errors and make mypy and E2E blocking in CI.
- Redeploy, if wanted, following docs/DEPLOY.md.
MIT. See LICENSE.
Pyae Sone (Seon) · github.com/soneeee22000




