Skip to content

proposal: use Jev via Vercel AI Gateway for bounded review routing #65

Description

@EthanThatOneKid

Summary

Evaluate Jev through Vercel AI Gateway as a bounded decision layer for Computer's PR-review pipeline. Jev should classify and route work, score risk, and gate expensive verification or human review; it should not replace the generative reviewer, implementer, or human approval boundary.

This is a follow-up to #57 and #20. It should only be promoted after the first reliable real run has produced a retrievable run record and real draft PR.

Why this fits Computer

Jev is a System One evaluation model. It evaluates shared state against typed questions and returns structured choices, scores, boolean probabilities, and confidence rather than prose. That makes it a good fit for narrow decisions such as:

  • documentation-only vs. standard feature vs. high-risk change;
  • whether a diff touches public API, authentication, permissions, migrations, or other sensitive paths;
  • whether to run deeper probes and validation;
  • whether evidence is sufficient to route to a generative reviewer or a person;
  • whether a review result is eligible for the next gate, subject to existing human approval rules.

Jev must not be used to write inline review comments, explain failures, suggest code fixes, or make an unconditional merge decision. Those remain responsibilities of the generative reviewer, CI, and humans.

Proposed experiment

  1. Add an opt-in Jev adapter behind a feature flag and a hard per-run/token budget.
  2. Send only the minimum redacted state required: diff metadata or bounded diff, touched paths, issue classification, CI/check results, and existing review evidence. Do not send secrets, credentials, raw chat, or untrusted issue text as instructions.
  3. Evaluate a fixed question set in one request and record the model ID, question IDs, typed answers, confidence/probability metadata, threshold version, and decision in durable run history.
  4. Route light, standard, and deep review paths from Jev's result, with uncertainty and high-risk signals always escalating rather than bypassing review.
  5. Run in shadow mode against historical or already-reviewed PRs first. Compare Jev routing with the established policy and measure false-light and false-deep decisions before allowing it to influence live runs.
  6. Keep the existing policy as the fail-closed fallback when Jev is unavailable, over budget, below confidence threshold, or returns an unexpected answer.

Acceptance criteria

  • Jev is called through the documented Vercel AI Gateway / AI SDK evaluation interface, not an ad-hoc provider endpoint.
  • The adapter has explicit opt-in, a bounded budget, timeout, retry limit, and fail-closed behavior.
  • The first phase is shadow-only and does not change PR labels, approvals, merge state, or review depth.
  • A calibration fixture covers documentation, ordinary feature, public API, security/auth, permissions, migration, and ambiguous changes.
  • Thresholds are evaluated against labeled examples and the cost of false-light decisions is documented.
  • Untrusted issue/PR text cannot alter the question definitions, policy thresholds, repository scope, or tool permissions.
  • Run history records the redacted input fingerprint, typed result, confidence, policy version, fallback path, and final routing decision.
  • The system never treats Jev's result alone as approval to merge or publish.
  • Cost, latency, accuracy, escalation rate, and provider availability are reported before expanding beyond shadow mode.

Pricing and availability constraint

Do not assume that Jev is unlimited or free. The current official Vercel AI Gateway model page lists Jev at $0.04 per 1M input tokens, with no output-token charge shown; the page does not show an unlimited free tier. The implementation must therefore use budgets and usage monitoring, and must verify current pricing and terms before enabling it in unattended runs.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestwayfinder:taskWayfinder task ticket (unblocks a decision)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions