Evaluation discipline for agentic apps — a vendored, local-first LLM-evaluation framework operated through your coding agent, with typed authority and a single Verdict Engine that gates only what you approve. The pure core; you decide what to measure and it derives no metrics on its own.
python calibration baseline scorecard codex ai-safety evaluation-framework verdict local-first llm agentic evals llm-evaluation claude-code ci-gating
-
Updated
Sep 14, 2026 - Python