Skip to content

Latest commit

 

History

30 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llmgate

CI Go Reference Go Report Card License: Apache 2.0

LiteLLM for Go. A single provider-agnostic client over Anthropic, OpenAI, Gemini, and the rest — with a router, fallback, cost tracking, and capability flags on top.

llmgate is the LLM gateway that vectorless-engine depends on, extracted into its own module so anything written in Go can use it. It is not a rewrite of LiteLLM. It sits on top of tmc/langchaingo — langchaingo handles every provider's wire protocol, llmgate wraps that behind a tiny Client interface and adds the production features langchaingo deliberately doesn't include.

How it fits together

architecture

The caller holds a Client. Middlewares (retry.New, budget.New, cache.New) wrap the client; a router.New can sit anywhere in the chain to fall over between providers. The private internal/adapter is the single seam where llmgate's interface meets langchaingo's provider implementations — one adapter serves all three providers.

middleware stack

Order matters. The outermost wrapper sees the call first; the innermost hits the network. Put cache.New below budget.New so cache hits don't burn budget. Put retry.New on top so retries run regardless of which inner layer tripped. Put limit.Client (or limit.Judge) just inside retry, so each attempt takes a slot and the 429 that triggers a retry has already narrowed the limit before the retry sleeps.

request flow

A call threads through the stack, optionally falls over to a backup provider, and comes back with Usage.{InputTokens, OutputTokens, TotalTokens, CostUSD} populated — no extra call, computed from a static price table.

Why this exists

The Go ecosystem has two extremes:

  • Vendor SDKs (openai-go, anthropic-sdk-go) — great typing, no portability. Swap providers, rewrite your call site.
  • Thin wrappers — portable, but you lose cost, retries, fallbacks, and capability introspection.

llmgate is the middle layer. One interface. Every provider behind it. All the production concerns — router, fallback on rate-limit, cost per call, capability flags — baked in rather than bolted on.

Status

Early code. Interface is the stable surface; implementations evolve underneath. Roadmap lives in ROADMAP.md.

What's in now:

  • Client interface with Complete + CountTokens

  • Anthropic, OpenAI, Gemini — all backed by langchaingo/llms, in the provider/ subpackages

  • A single internal adapter; add a provider = add a ~30-line file

  • retry.New middleware for exp-backoff on transient errors

  • limit.New — an adaptive per-provider concurrency limiter (halves on a 429 or transport failure, honours Retry-After, widens on sustained success); wraps Client and Judge, and reports every change through OnChange

  • Cost tracking via a static pricing table + Usage.CostUSD on every Response

  • Capability flags (MaxContext, SupportsJSONMode, SupportsStreaming, SupportsTools, SupportsVision) with a capabilities.Capable interface

  • router.New with per-provider fallback (router.OnRateLimit, router.OnTransient, or a custom router.FallbackPolicy)

  • budget.New middleware with daily + total USD caps and UTC rollover

  • cache.New middleware — in-memory LRU keyed on request shape, optional TTL

  • Error classification (Classify, IsRateLimited, IsTransient, IsAuth) that the retry predicate and router policies use to decide what to retry or fall over on

  • Mock client with call recording for tests

  • Native tool calling (Request.Tools, Response.ToolCalls) across all three providers

Coming next:

  • Streaming (Streamer is declared; no provider implements it yet)
  • Native count_tokens via each provider's counting endpoint
  • Anthropic prompt caching and OpenAI strict structured outputs, both of which need a native client rather than the langchaingo path

Install

go get github.com/hallelx2/llmgate

Use

package main

import (
    "context"
    "fmt"
    "os"

    "github.com/hallelx2/llmgate"
    "github.com/hallelx2/llmgate/middleware/retry"
    "github.com/hallelx2/llmgate/provider/anthropic"
)

func main() {
    client, err := anthropic.New(anthropic.Config{
        APIKey: os.Getenv("ANTHROPIC_API_KEY"),
        Model:  "claude-sonnet-4-5",
    })
    if err != nil {
        panic(err)
    }

    // Optional: wrap in exponential-backoff retry middleware.
    client = retry.New(retry.Config{MaxRetries: 3})(client)

    resp, err := client.Complete(context.Background(), llmgate.Request{
        Messages: []llmgate.Message{
            {Role: llmgate.RoleUser, Content: "In one sentence: what is vectorless retrieval?"},
        },
        MaxTokens: 256,
    })
    if err != nil {
        panic(err)
    }

    fmt.Println(resp.Content)
}

System One: judgments instead of text

Some work does not need language. Routing a ticket, ranking a candidate, checking whether a claim holds — the output is a decision, and wrapping it in prose only to parse the prose back out is pure overhead.

Judge is a second interface for models that return that decision directly. It sits beside Client, not behind it: a chat-completion surface has roles, messages and an assistant turn, and a System One model has none of those. Hold a Client for work that must produce language, a Judge for work that must produce a decision, and both when you need both.

The first implementation is TypeSafe's Jev.

j, _ := typesafe.New(typesafe.Config{}) // reads TYPESAFE_API_KEY
judge := retry.NewJudge(retry.Config{MaxRetries: 3})(j)

result, err := judge.Judge(ctx, llmgate.JudgeRequest{
    Model: "jev-1.13.0",
    State: map[string]any{"ticket": ticket},
    Questions: map[string]llmgate.Question{
        "is_urgent":    llmgate.Noul{Instructions: "Is this time-sensitive?"},
        "wants_refund": llmgate.Noul{Instructions: "Is the customer asking for money back?"},
        "team": llmgate.Choice{
            Instructions: "Which team should own this?",
            Options: llmgate.ChoiceOptions{
                {Name: "billing", Description: "Payments, invoicing, refunds"},
                {Name: "technical", Description: "Bugs, outages, integrations"},
                {Name: "unclear", Description: "Not enough information to route"},
            },
        },
        "frustration": llmgate.Score{
            Instructions: "How frustrated does the customer sound?",
            Levels:       []string{"Neutral", "Annoyed", "Angry", "Threatening to leave"},
        },
    },
})

urgent, _ := result.Noul("is_urgent")      // 0.92
team, _ := result.Choice("team")           // .Choice, .Probabilities, .Confidence
mood, _ := result.Score("frustration")     // .Score, .Legend, .Probabilities, .Confidence

Runnable version: examples/judge.

Live tests sit behind a build tag so a plain go test ./... stays offline and free:

# key from the environment, or from a gitignored .env up-tree
go test -tags live -v -run TestLive ./judge/typesafe/

TestLiveLatencyScaling is the one worth re-running when the model version moves — it measures whether a batch of N questions still costs roughly what one costs.

One request, many questions

That example is one HTTP call, not four. The model reads the state once and runs every question against it in parallel, so packing a batch is the whole point — the instinct to fan out into a call per question is the thing to unlearn. Questions in one request cannot see each other's answers, which is exactly what makes them parallelisable; send a second request only when an answer decides what evidence to fetch next.

Two limits bound a batch: 64k tokens for the state plus all questions, and 32k for the state plus the single longest question. The transport checks both before sending, so an oversized batch costs nothing.

Three primitives

Returns Reach for it when
Noul probability of yes A condition either holds or it doesn't. Use one per label when several can apply at once
Choice one option + the full distribution + confidence Exactly one outcome wins
Score probability-weighted position + distribution + confidence The answer is a degree along an ordered rubric

Noul carries no confidence, and that is not an omission — the probability is the answer. A Noul near 0.5 means yes and no are about equally likely; it does not mean "medium intensity".

Confidence on the other two summarises how concentrated the distribution is. It is not a probability that the answer is correct, and it is not permission to act. Several equally acceptable options also spread probability, so low confidence on a harmless choice is not a problem to route around.

What it is bad at

TypeSafe publishes a jaggedness page for Jev. The three that will bite hardest:

  • It is not a calculator. Counting, arithmetic, and date ordering are all unreliable — it reads dates as text, not as ordered quantities. Keep every one of those in code. Extraction is a judgment; comparison is not.
  • It reads literally. It answers the question you wrote, not the one you meant. Keep Instructions and the criteria aligned — a Noul whose true describes the "no" case measurably underperforms. When you find yourself explaining what you really meant, that explanation is the missing half of the instruction.
  • Context rot. Accuracy falls as the state grows with material unrelated to the decision. Filter first, and send only the fields the questions need.

It also does not generate text. When the answer space is bounded, turn extraction into a Choice over candidates your code found, rather than asking for the value itself.

Cost

Jev bills input tokens only; output is free. Rates live in pricing/systemone.go — hand-maintained, because the upstream price feeds only index chat-completion models and a generated entry would be wiped by the next go generate. It is consulted last, so a feed or an explicit pricing.Register still wins.

Cost accounting

Every Response carries a Usage with the token breakdown split by billing tier, not just a prompt/completion pair:

resp.Usage.InputTokens      // uncached prompt tokens
resp.Usage.CacheWriteTokens // written to the prompt cache (1.25x on Anthropic)
resp.Usage.CacheReadTokens  // served from cache (0.1x Anthropic, 0.5x OpenAI, 0.25x Google)
resp.Usage.ReasoningTokens  // thinking tokens — a subset of OutputTokens
resp.Usage.CostUSD

The providers disagree about whether cached tokens are counted inside the prompt total — Anthropic reports them alongside, OpenAI and Google report them within — so llmgate normalizes to the disjoint form. InputTokens + CacheWriteTokens + CacheReadTokens is the whole prompt on every provider.

Three flags say how much to trust the number:

Flag Meaning when false
Priced no price-book entry — CostUSD is unknown, not zero
TokensReported the provider returned no counts
Estimated (when true) counts came from a local tokenizer, so cost is approximate

That distinction matters: a CostUSD of 0 with Priced true asserts the call was free, which it never is. When a provider reports no usage at all, llmgate estimates from a tokenizer and flags it rather than reporting zero.

Model IDs are normalized

Lookups resolve dated, prefixed, and gateway-qualified IDs to their base model, so none of these price at $0:

claude-sonnet-4-5-20250929        -> claude-sonnet-4-5
models/gemini-2.5-flash           -> gemini-2.5-flash
us.anthropic.claude-opus-4-1-v1:0 -> claude-opus-4-1
z-ai/glm-4.6                      -> glm-4.6

Rates are keyed by model ID alone, never by provider — a model served through a gateway that speaks another vendor's protocol (GLM over an Anthropic-compatible endpoint, say) still prices correctly.

Live prices (opt-in)

The compiled-in table drifts as vendors change rates. No vendor publishes machine-readable pricing, so UseRemote layers a community aggregate over it:

stop, err := pricing.UseRemote(ctx, pricing.RemoteConfig{
    CacheDir: "/var/cache/llmgate", // survive restarts
    OnError:  func(src string, err error) { log.Warn("price refresh", "src", src, "err", err) },
})
defer stop()

Sources default to LiteLLM's price table then OpenRouter's model API, both public and unauthenticated.

This is opt-in and stays that way — importing pricing performs no network I/O. Resolution is Register overrides → remote snapshot → embedded table, and it fails open at every layer: lookups never block on the network, a failed fetch keeps the previous snapshot, and a snapshot whose rates have drifted more than 10x from the embedded values is rejected as corrupt rather than adopted. pricing.AsOf() reports the vintage of whatever is loaded.

Design principles

  • The interface is tiny. Complete, CountTokens, later Stream, later Capabilities. If it doesn't fit, it goes in a middleware, not the interface.
  • A second interface beats a leaky first one. Judge exists because a System One model has no messages and no assistant turn. Forcing it through Complete would mean encoding questions into a prompt and parsing answers back out of text — reintroducing exactly the JSON-mode retry machinery that a typed API removes. When a model's shape genuinely differs, add a seam; do not widen the existing one until it fits everything and describes nothing.
  • Middleware over inheritance. Retries, caching, cost tracking, rate-limiting — all func(Client) Client wrappers. Compose them.
  • No magic config. No viper, no auto-reload, no remote backends. Pass a struct, get a client.
  • Provider-specific features are honoured where they matter. Anthropic prompt caching, OpenAI structured outputs, Gemini long context — opt in via config, not via a lowest-common-denominator API.
  • Pure IO-bound code. Parallelism is always network-bound. errgroup + semaphore, no worker pools.

See ROADMAP.md for what is shipped and what is next.

License

Apache 2.0. See LICENSE.

About

LiteLLM for Go. Provider-agnostic LLM client over Anthropic, OpenAI, Gemini with router, fallback, cost tracking, capability flags, and composable middleware — built on tmc/langchaingo.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages