LiteLLM for Go. A single provider-agnostic client over Anthropic, OpenAI, Gemini, and the rest — with a router, fallback, cost tracking, and capability flags on top.
llmgate is the LLM gateway that vectorless-engine
depends on, extracted into its own module so anything written in Go
can use it. It is not a rewrite of LiteLLM. It sits on top of
tmc/langchaingo — langchaingo
handles every provider's wire protocol, llmgate wraps that behind a
tiny Client interface and adds the production features langchaingo
deliberately doesn't include.
The caller holds a Client. Middlewares (retry.New, budget.New,
cache.New) wrap the client; a router.New can sit anywhere in the chain
to fall over between providers. The private internal/adapter is the
single seam where llmgate's interface meets langchaingo's provider
implementations — one adapter serves all three providers.
Order matters. The outermost wrapper sees the call first; the innermost
hits the network. Put cache.New below budget.New so cache hits
don't burn budget. Put retry.New on top so retries run regardless
of which inner layer tripped. Put limit.Client (or limit.Judge)
just inside retry, so each attempt takes a slot and the 429 that
triggers a retry has already narrowed the limit before the retry
sleeps.
A call threads through the stack, optionally falls over to a backup
provider, and comes back with Usage.{InputTokens, OutputTokens, TotalTokens, CostUSD} populated — no extra call, computed from a
static price table.
The Go ecosystem has two extremes:
- Vendor SDKs (
openai-go,anthropic-sdk-go) — great typing, no portability. Swap providers, rewrite your call site. - Thin wrappers — portable, but you lose cost, retries, fallbacks, and capability introspection.
llmgate is the middle layer. One interface. Every provider behind
it. All the production concerns — router, fallback on rate-limit,
cost per call, capability flags — baked in rather than bolted on.
Early code. Interface is the stable surface; implementations evolve underneath. Roadmap lives in ROADMAP.md.
What's in now:
-
Clientinterface withComplete+CountTokens -
Anthropic, OpenAI, Gemini — all backed by
langchaingo/llms, in theprovider/subpackages -
A single internal adapter; add a provider = add a ~30-line file
-
retry.Newmiddleware for exp-backoff on transient errors -
limit.New— an adaptive per-provider concurrency limiter (halves on a 429 or transport failure, honoursRetry-After, widens on sustained success); wrapsClientandJudge, and reports every change throughOnChange -
Cost tracking via a static
pricingtable +Usage.CostUSDon every Response -
Capability flags (
MaxContext,SupportsJSONMode,SupportsStreaming,SupportsTools,SupportsVision) with acapabilities.Capableinterface -
router.Newwith per-provider fallback (router.OnRateLimit,router.OnTransient, or a customrouter.FallbackPolicy) -
budget.Newmiddleware with daily + total USD caps and UTC rollover -
cache.Newmiddleware — in-memory LRU keyed on request shape, optional TTL -
Error classification (
Classify,IsRateLimited,IsTransient,IsAuth) that the retry predicate and router policies use to decide what to retry or fall over on -
Mockclient with call recording for tests -
Native tool calling (
Request.Tools,Response.ToolCalls) across all three providers
Coming next:
- Streaming (
Streameris declared; no provider implements it yet) - Native
count_tokensvia each provider's counting endpoint - Anthropic prompt caching and OpenAI strict structured outputs, both of which need a native client rather than the langchaingo path
go get github.com/hallelx2/llmgatepackage main
import (
"context"
"fmt"
"os"
"github.com/hallelx2/llmgate"
"github.com/hallelx2/llmgate/middleware/retry"
"github.com/hallelx2/llmgate/provider/anthropic"
)
func main() {
client, err := anthropic.New(anthropic.Config{
APIKey: os.Getenv("ANTHROPIC_API_KEY"),
Model: "claude-sonnet-4-5",
})
if err != nil {
panic(err)
}
// Optional: wrap in exponential-backoff retry middleware.
client = retry.New(retry.Config{MaxRetries: 3})(client)
resp, err := client.Complete(context.Background(), llmgate.Request{
Messages: []llmgate.Message{
{Role: llmgate.RoleUser, Content: "In one sentence: what is vectorless retrieval?"},
},
MaxTokens: 256,
})
if err != nil {
panic(err)
}
fmt.Println(resp.Content)
}Some work does not need language. Routing a ticket, ranking a candidate, checking whether a claim holds — the output is a decision, and wrapping it in prose only to parse the prose back out is pure overhead.
Judge is a second interface for models that return that decision
directly. It sits beside Client, not behind it: a chat-completion
surface has roles, messages and an assistant turn, and a System One model
has none of those. Hold a Client for work that must produce language, a
Judge for work that must produce a decision, and both when you need
both.
The first implementation is TypeSafe's Jev.
j, _ := typesafe.New(typesafe.Config{}) // reads TYPESAFE_API_KEY
judge := retry.NewJudge(retry.Config{MaxRetries: 3})(j)
result, err := judge.Judge(ctx, llmgate.JudgeRequest{
Model: "jev-1.13.0",
State: map[string]any{"ticket": ticket},
Questions: map[string]llmgate.Question{
"is_urgent": llmgate.Noul{Instructions: "Is this time-sensitive?"},
"wants_refund": llmgate.Noul{Instructions: "Is the customer asking for money back?"},
"team": llmgate.Choice{
Instructions: "Which team should own this?",
Options: llmgate.ChoiceOptions{
{Name: "billing", Description: "Payments, invoicing, refunds"},
{Name: "technical", Description: "Bugs, outages, integrations"},
{Name: "unclear", Description: "Not enough information to route"},
},
},
"frustration": llmgate.Score{
Instructions: "How frustrated does the customer sound?",
Levels: []string{"Neutral", "Annoyed", "Angry", "Threatening to leave"},
},
},
})
urgent, _ := result.Noul("is_urgent") // 0.92
team, _ := result.Choice("team") // .Choice, .Probabilities, .Confidence
mood, _ := result.Score("frustration") // .Score, .Legend, .Probabilities, .ConfidenceRunnable version: examples/judge.
Live tests sit behind a build tag so a plain go test ./... stays offline
and free:
# key from the environment, or from a gitignored .env up-tree
go test -tags live -v -run TestLive ./judge/typesafe/TestLiveLatencyScaling is the one worth re-running when the model version
moves — it measures whether a batch of N questions still costs roughly what
one costs.
That example is one HTTP call, not four. The model reads the state once and runs every question against it in parallel, so packing a batch is the whole point — the instinct to fan out into a call per question is the thing to unlearn. Questions in one request cannot see each other's answers, which is exactly what makes them parallelisable; send a second request only when an answer decides what evidence to fetch next.
Two limits bound a batch: 64k tokens for the state plus all questions, and 32k for the state plus the single longest question. The transport checks both before sending, so an oversized batch costs nothing.
| Returns | Reach for it when | |
|---|---|---|
Noul |
probability of yes | A condition either holds or it doesn't. Use one per label when several can apply at once |
Choice |
one option + the full distribution + confidence | Exactly one outcome wins |
Score |
probability-weighted position + distribution + confidence | The answer is a degree along an ordered rubric |
Noul carries no confidence, and that is not an omission — the
probability is the answer. A Noul near 0.5 means yes and no are about
equally likely; it does not mean "medium intensity".
Confidence on the other two summarises how concentrated the
distribution is. It is not a probability that the answer is correct, and
it is not permission to act. Several equally acceptable options also
spread probability, so low confidence on a harmless choice is not a
problem to route around.
TypeSafe publishes a jaggedness page for Jev. The three that will bite hardest:
- It is not a calculator. Counting, arithmetic, and date ordering are all unreliable — it reads dates as text, not as ordered quantities. Keep every one of those in code. Extraction is a judgment; comparison is not.
- It reads literally. It answers the question you wrote, not the one
you meant. Keep
Instructionsand the criteria aligned — aNoulwhosetruedescribes the "no" case measurably underperforms. When you find yourself explaining what you really meant, that explanation is the missing half of the instruction. - Context rot. Accuracy falls as the state grows with material unrelated to the decision. Filter first, and send only the fields the questions need.
It also does not generate text. When the answer space is bounded, turn
extraction into a Choice over candidates your code found, rather than
asking for the value itself.
Jev bills input tokens only; output is free. Rates live in
pricing/systemone.go — hand-maintained, because the upstream price feeds
only index chat-completion models and a generated entry would be wiped by
the next go generate. It is consulted last, so a feed or an explicit
pricing.Register still wins.
Every Response carries a Usage with the token breakdown split by
billing tier, not just a prompt/completion pair:
resp.Usage.InputTokens // uncached prompt tokens
resp.Usage.CacheWriteTokens // written to the prompt cache (1.25x on Anthropic)
resp.Usage.CacheReadTokens // served from cache (0.1x Anthropic, 0.5x OpenAI, 0.25x Google)
resp.Usage.ReasoningTokens // thinking tokens — a subset of OutputTokens
resp.Usage.CostUSDThe providers disagree about whether cached tokens are counted inside the
prompt total — Anthropic reports them alongside, OpenAI and Google report
them within — so llmgate normalizes to the disjoint form. InputTokens + CacheWriteTokens + CacheReadTokens is the whole prompt on every provider.
Three flags say how much to trust the number:
| Flag | Meaning when false |
|---|---|
Priced |
no price-book entry — CostUSD is unknown, not zero |
TokensReported |
the provider returned no counts |
Estimated |
(when true) counts came from a local tokenizer, so cost is approximate |
That distinction matters: a CostUSD of 0 with Priced true asserts the
call was free, which it never is. When a provider reports no usage at all,
llmgate estimates from a tokenizer and flags it rather than reporting zero.
Lookups resolve dated, prefixed, and gateway-qualified IDs to their base model, so none of these price at $0:
claude-sonnet-4-5-20250929 -> claude-sonnet-4-5
models/gemini-2.5-flash -> gemini-2.5-flash
us.anthropic.claude-opus-4-1-v1:0 -> claude-opus-4-1
z-ai/glm-4.6 -> glm-4.6
Rates are keyed by model ID alone, never by provider — a model served through a gateway that speaks another vendor's protocol (GLM over an Anthropic-compatible endpoint, say) still prices correctly.
The compiled-in table drifts as vendors change rates. No vendor publishes
machine-readable pricing, so UseRemote layers a community aggregate over
it:
stop, err := pricing.UseRemote(ctx, pricing.RemoteConfig{
CacheDir: "/var/cache/llmgate", // survive restarts
OnError: func(src string, err error) { log.Warn("price refresh", "src", src, "err", err) },
})
defer stop()Sources default to LiteLLM's price table then OpenRouter's model API, both public and unauthenticated.
This is opt-in and stays that way — importing pricing performs no
network I/O. Resolution is Register overrides → remote snapshot →
embedded table, and it fails open at every layer: lookups never block on
the network, a failed fetch keeps the previous snapshot, and a snapshot
whose rates have drifted more than 10x from the embedded values is
rejected as corrupt rather than adopted. pricing.AsOf() reports the
vintage of whatever is loaded.
- The interface is tiny.
Complete,CountTokens, laterStream, laterCapabilities. If it doesn't fit, it goes in a middleware, not the interface. - A second interface beats a leaky first one.
Judgeexists because a System One model has no messages and no assistant turn. Forcing it throughCompletewould mean encoding questions into a prompt and parsing answers back out of text — reintroducing exactly the JSON-mode retry machinery that a typed API removes. When a model's shape genuinely differs, add a seam; do not widen the existing one until it fits everything and describes nothing. - Middleware over inheritance. Retries, caching, cost tracking,
rate-limiting — all
func(Client) Clientwrappers. Compose them. - No magic config. No viper, no auto-reload, no remote backends. Pass a struct, get a client.
- Provider-specific features are honoured where they matter. Anthropic prompt caching, OpenAI structured outputs, Gemini long context — opt in via config, not via a lowest-common-denominator API.
- Pure IO-bound code. Parallelism is always network-bound.
errgroup+semaphore, no worker pools.
See ROADMAP.md for what is shipped and what is next.
Apache 2.0. See LICENSE.