Typed Go clients for OpenAI-compatible providers, Anthropic-compatible servers, native llama.cpp, and Ollama. The library calls provider HTTP and WebSocket APIs directly and exposes provider-native models organized by bounded context.
- OpenAI-compatible REST client with custom base URLs and per-request auth.
- OpenAI Codex device-code login with refreshable bearer sessions.
- OpenAI domain packages for responses, chat, completions, embeddings, models, files, uploads, batches, media, stateful resources, organization APIs, and Realtime WebSocket sessions.
- Ollama client with generation, chat, embeddings, model management, blobs, and version APIs.
- Anthropic Messages, Models, Files, and Message Batches clients, including typed SSE events and cursor pagination.
- Native llama.cpp client for completion, embeddings, tokenization, templates, server properties, slots, LoRA, metrics, router models, and reranking.
- Typed SSE and NDJSON iterators with explicit
Next,Value,Err, andCloselifecycle methods. - Provider-independent environment metadata through
ollama.EnvironmentScheme()andllama.EnvironmentScheme(); each entry includes the variable name, finite allowed values when applicable, default, and operational description. - Bounded response and event sizes, context cancellation, typed protocol errors,
and deterministic
httptest-based tests. easyjson-generated provider models;github.com/coder/websocketis used only by the OpenAI Realtime package.
go get go.osspkg.com/llm-clientThe module requires Go 1.26.8 or newer. See the capability matrix for endpoint-level coverage and the API reference for package details.
The following example creates an OpenAI-compatible client and sends a chat
request. Set OPENAI_API_KEY before running it.
package main
import (
"context"
"encoding/json"
"fmt"
"os"
"go.osspkg.com/llm-client/openai"
"go.osspkg.com/llm-client/openai/chat"
"go.osspkg.com/llm-client/pkg/auth"
)
func main() {
client, err := openai.New(
openai.WithAuthProvider(auth.StaticBearer(os.Getenv("OPENAI_API_KEY"))),
)
if err != nil {
panic(err)
}
response, err := client.Chat().Create(context.Background(), chat.Request{
Model: "model",
Messages: []chat.Message{{
Role: "user",
Content: json.RawMessage([]byte(`"Hello"`)),
}},
})
if err != nil {
panic(err)
}
fmt.Println(string(response.Choices[0].Message.Content))
}For an OpenAI-compatible service, provide its endpoint explicitly:
client, err := openai.New(
openai.WithBaseURL("http://localhost:8080/v1"),
openai.WithAuthProvider(auth.StaticBearer("token")),
)For Codex subscription authentication, use the device-code flow. The callback shows the short-lived code in the browser or terminal, and the session refreshes access tokens when the provider supplies an expiry:
codexAuth, err := auth.NewCodexDeviceAuth()
if err != nil {
panic(err)
}
codexSession, err := codexAuth.Login(ctx, func(code auth.CodexDeviceCode) error {
fmt.Printf("Open %s and enter %s\n", code.VerificationURL, code.UserCode)
return nil
})
if err != nil {
panic(err)
}
client, err := openai.New(openai.WithAuthProvider(codexSession.HeaderProvider()))Anthropic uses x-api-key and the required API version by default:
client, err := anthropic.New(anthropic.WithAPIKey(os.Getenv("ANTHROPIC_API_KEY")))
content, err := messages.TextContent("Hello")
if err != nil {
panic(err)
}
response, err := client.Messages().Create(ctx, messages.Request{
Model: "claude-3-5-sonnet-latest", MaxTokens: 256,
Messages: []messages.Message{{Role: "user", Content: content}},
})For a native llama.cpp server, use the separate native client. Its default
endpoint is http://localhost:8080:
client, err := llama.New()
prompt, err := completions.StringPrompt("Write one short sentence.")
if err != nil {
panic(err)
}
response, err := client.Completions().Create(ctx, completions.Request{
Prompt: prompt,
NPredict: 32,
})Ollama uses http://localhost:11434 by default:
package main
import (
"context"
"fmt"
"go.osspkg.com/llm-client/ollama"
"go.osspkg.com/llm-client/ollama/chat"
)
func main() {
client, err := ollama.New()
if err != nil {
panic(err)
}
response, err := client.Chat().Create(context.Background(), chat.Request{
Model: "llama3.2",
Messages: []chat.Message{{Role: "user", Content: "Hello"}},
})
if err != nil {
panic(err)
}
fmt.Println(response.Message.Content)
}Streaming calls return typed iterators. The caller owns the iterator and must close it when processing is complete.
events, err := client.Chat().CreateStream(ctx, chat.Request{
Model: "model",
Messages: []chat.Message{{
Role: "user",
Content: json.RawMessage([]byte(`"Tell me a short joke."`)),
}},
})
if err != nil {
return err
}
defer events.Close()
for events.Next(ctx) {
chunk := events.Value()
// Process the typed chunk.
_ = chunk
}
if err := events.Err(); err != nil {
return err
}The same iterator contract is used for OpenAI SSE and Ollama NDJSON streams.
See pkg/stream and the streaming sections in DOC.md
for parser limits and error handling.
The Ollama and native llama.cpp clients expose server configuration metadata without requiring a client instance:
for _, variable := range ollama.EnvironmentScheme().Variables {
fmt.Printf("%s (default %q): %s\n", variable.Name, variable.Default, variable.Description)
}AllowedValues is populated for finite enumerations. An empty slice means the
provider accepts a documented scalar format such as a path, duration, integer,
or comma-separated list; the required format is described in Description.
The function returns metadata only and does not read or modify the process
environment. Use llama.EnvironmentScheme() for llama-server variables.
| Provider | Package | Transport | Default endpoint |
|---|---|---|---|
| OpenAI-compatible | openai |
HTTP; Realtime WebSocket | https://api.openai.com/v1 |
| Anthropic-compatible | anthropic |
HTTP; SSE | https://api.anthropic.com/v1 |
| Native llama.cpp | llama |
HTTP; SSE | http://localhost:8080 |
| Ollama | ollama |
HTTP; NDJSON streaming | http://localhost:11434 |
Provider APIs remain separate. There is no provider-neutral facade that hides provider-specific capabilities or request models.
Root clients are immutable after construction and safe for concurrent use;
domain clients are obtained through accessors such as client.Chat() and
client.Completions().
For the current endpoint and capability status, see:
docs/capability-matrix.mdDOC.md— English API and architecture referenceDOC.ru.md— Russian guide
Generated easyjson files are committed to the repository and must not be
edited manually. Run the complete local quality workflow:
go generate ./...
make lint
make tests
go test -race ./...
go vet ./...
go mod verify
git diff --checkTests use deterministic mocks and do not require provider credentials or live
network access. See AGENTS.md for architecture, dependency,
serialization, security, and resource-lifecycle rules.
Contributions are welcome. Before opening a pull request:
- Keep provider-specific types and operations in their provider/domain package.
- Keep shared transport, auth, error, pagination, stream, and WebSocket
lifecycle code in
pkgwithout provider imports. - Regenerate easyjson output with
go generate ./.... - Run the development checks listed above.
- Add deterministic tests for new protocol behavior and resource ownership.
This project is available under the BSD 3-Clause License.