Ship evals before you ship features.
-
Updated
Jun 22, 2026 - Nunjucks
Ship evals before you ship features.
Eval framework. Define correct, test against it, get results.
A guard-railed, closed-loop workflow for AI coding agents: live state bus + execution-level hard intercepts for Claude Code, Codex, and DeepSeek Harness (GitHub PR / GitLab MR). From step-level to requirement-level; eval-driven, spec-driven, human-in-the-loop.
AI-augmented QA platform for spec-driven development and testing, RAG-grounded analysis, eval-driven development and contract validation across Python, Go, Rust and Solidity.
Autonomous skill improvement loop for Claude Code plugins — inspired by Karpathy's autoresearch. Modify → evaluate → keep/discard → repeat until convergence. Zero-touch quality iteration at scale.
NeuroSwarm — the agentic layer harness: one protocol, 18 agents, 15 commands, 8 rules, 13 skills, a 19-pattern NestJS stack and native adapters for opencode, Claude Code, Cursor, Copilot and Gemini. Switch-controlled token loading, eval-driven development, feature tracking and measurable metrics.
正解表なしで5つの性質からソートを採点するメタモルフィック・オラクル | Sort graded by metamorphic relations
非決定的なシャッフルをカイ二乗検定で採点する統計オラクル | Shuffle graded by Monte-Carlo statistical tests
LLM査読の検出力をラベル付き見本で採点するメタ評価ハーネス | Meta-eval: grading an LLM reviewer with labeled specimens
Production harness for a multi-agent BI system — eval-gated, guardrailed, cross-source-validated. LangGraph + hybrid RAG + FastAPI, live on AWS.
Multi-agent inspection pipeline for solar cell EL images: EfficientNet-B0 severity classifier + Qwen3-VL (Ollama) reasoning, served via FastAPI. 75.3% on a 20-criteria eval suite.
Companion code for the talk "Managing Production Agents at Scale — from Chaos to Reliability". One Google ADK agent, three production failure modes: eval-driven development, resilience, and zero-trust on Vertex AI.
Claude Codeエージェント定義の査読と、その検出力を採点するオラクル | Agent-spec review graded by a detection-power oracle
A hands-on learning repository exploring Spec-Driven Development (SDD) for building deterministic AI systems. Covers specs, evaluation loops, patterns, experiments, and failures to bridge theory with real-world AI engineering practices.
ランダム入力の集中砲火で「壊れない」を採点するファジング・オラクル | Robust parser graded by a fuzzing (implicit) oracle
Fractal design docs: one frame, zoom into any element, same shape at every depth — with a document-structure oracle | 枠1枚から掘れるフラクタル設計書の生成エージェント+文書構造オラクル
答えが一意でない出力を仕様アサーションで採点するEDD実証 | Divisor finder graded by a spec-assertion oracle
Eval 驱动的 LLM Agent 应用开发团队(Claude Code Subagents)——没有 eval 不许改 prompt
Eval-driven development for LLM accounting skills. 50 test cases · 66% → 100% in 6 iterations · results reproducible with the included grader
Most AI plugins hope they work. These prove it. Eval-driven Claude plugins for product teams.
To associate your repository with the eval-driven-development topic, visit your repo's landing page and select "manage topics."