Skip to content

Repository files navigation

go-llm-client

Go Reference Go Report Card Go version License

Typed Go clients for OpenAI-compatible providers, Anthropic-compatible servers, native llama.cpp, and Ollama. The library calls provider HTTP and WebSocket APIs directly and exposes provider-native models organized by bounded context.

Contents

Features

  • OpenAI-compatible REST client with custom base URLs and per-request auth.
  • OpenAI Codex device-code login with refreshable bearer sessions.
  • OpenAI domain packages for responses, chat, completions, embeddings, models, files, uploads, batches, media, stateful resources, organization APIs, and Realtime WebSocket sessions.
  • Ollama client with generation, chat, embeddings, model management, blobs, and version APIs.
  • Anthropic Messages, Models, Files, and Message Batches clients, including typed SSE events and cursor pagination.
  • Native llama.cpp client for completion, embeddings, tokenization, templates, server properties, slots, LoRA, metrics, router models, and reranking.
  • Typed SSE and NDJSON iterators with explicit Next, Value, Err, and Close lifecycle methods.
  • Provider-independent environment metadata through ollama.EnvironmentScheme() and llama.EnvironmentScheme(); each entry includes the variable name, finite allowed values when applicable, default, and operational description.
  • Bounded response and event sizes, context cancellation, typed protocol errors, and deterministic httptest-based tests.
  • easyjson-generated provider models; github.com/coder/websocket is used only by the OpenAI Realtime package.

Installation

go get go.osspkg.com/llm-client

The module requires Go 1.26.8 or newer. See the capability matrix for endpoint-level coverage and the API reference for package details.

Quick start

The following example creates an OpenAI-compatible client and sends a chat request. Set OPENAI_API_KEY before running it.

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"os"

	"go.osspkg.com/llm-client/openai"
	"go.osspkg.com/llm-client/openai/chat"
	"go.osspkg.com/llm-client/pkg/auth"
)

func main() {
	client, err := openai.New(
		openai.WithAuthProvider(auth.StaticBearer(os.Getenv("OPENAI_API_KEY"))),
	)
	if err != nil {
		panic(err)
	}

	response, err := client.Chat().Create(context.Background(), chat.Request{
		Model: "model",
		Messages: []chat.Message{{
			Role:    "user",
			Content: json.RawMessage([]byte(`"Hello"`)),
		}},
	})
	if err != nil {
		panic(err)
	}

	fmt.Println(string(response.Choices[0].Message.Content))
}

For an OpenAI-compatible service, provide its endpoint explicitly:

client, err := openai.New(
	openai.WithBaseURL("http://localhost:8080/v1"),
	openai.WithAuthProvider(auth.StaticBearer("token")),
)

For Codex subscription authentication, use the device-code flow. The callback shows the short-lived code in the browser or terminal, and the session refreshes access tokens when the provider supplies an expiry:

codexAuth, err := auth.NewCodexDeviceAuth()
if err != nil {
	panic(err)
}
codexSession, err := codexAuth.Login(ctx, func(code auth.CodexDeviceCode) error {
	fmt.Printf("Open %s and enter %s\n", code.VerificationURL, code.UserCode)
	return nil
})
if err != nil {
	panic(err)
}
client, err := openai.New(openai.WithAuthProvider(codexSession.HeaderProvider()))

Anthropic uses x-api-key and the required API version by default:

client, err := anthropic.New(anthropic.WithAPIKey(os.Getenv("ANTHROPIC_API_KEY")))
content, err := messages.TextContent("Hello")
if err != nil {
	panic(err)
}
response, err := client.Messages().Create(ctx, messages.Request{
	Model: "claude-3-5-sonnet-latest", MaxTokens: 256,
	Messages: []messages.Message{{Role: "user", Content: content}},
})

For a native llama.cpp server, use the separate native client. Its default endpoint is http://localhost:8080:

client, err := llama.New()
prompt, err := completions.StringPrompt("Write one short sentence.")
if err != nil {
	panic(err)
}
response, err := client.Completions().Create(ctx, completions.Request{
	Prompt: prompt,
	NPredict: 32,
})

Ollama uses http://localhost:11434 by default:

package main

import (
	"context"
	"fmt"

	"go.osspkg.com/llm-client/ollama"
	"go.osspkg.com/llm-client/ollama/chat"
)

func main() {
	client, err := ollama.New()
	if err != nil {
		panic(err)
	}

	response, err := client.Chat().Create(context.Background(), chat.Request{
		Model:    "llama3.2",
		Messages: []chat.Message{{Role: "user", Content: "Hello"}},
	})
	if err != nil {
		panic(err)
	}

	fmt.Println(response.Message.Content)
}

Streaming

Streaming calls return typed iterators. The caller owns the iterator and must close it when processing is complete.

events, err := client.Chat().CreateStream(ctx, chat.Request{
	Model: "model",
	Messages: []chat.Message{{
		Role:    "user",
		Content: json.RawMessage([]byte(`"Tell me a short joke."`)),
	}},
})
if err != nil {
	return err
}
defer events.Close()

for events.Next(ctx) {
	chunk := events.Value()
	// Process the typed chunk.
	_ = chunk
}
if err := events.Err(); err != nil {
	return err
}

The same iterator contract is used for OpenAI SSE and Ollama NDJSON streams. See pkg/stream and the streaming sections in DOC.md for parser limits and error handling.

Server environment schemes

The Ollama and native llama.cpp clients expose server configuration metadata without requiring a client instance:

for _, variable := range ollama.EnvironmentScheme().Variables {
	fmt.Printf("%s (default %q): %s\n", variable.Name, variable.Default, variable.Description)
}

AllowedValues is populated for finite enumerations. An empty slice means the provider accepts a documented scalar format such as a path, duration, integer, or comma-separated list; the required format is described in Description. The function returns metadata only and does not read or modify the process environment. Use llama.EnvironmentScheme() for llama-server variables.

Provider coverage

Provider Package Transport Default endpoint
OpenAI-compatible openai HTTP; Realtime WebSocket https://api.openai.com/v1
Anthropic-compatible anthropic HTTP; SSE https://api.anthropic.com/v1
Native llama.cpp llama HTTP; SSE http://localhost:8080
Ollama ollama HTTP; NDJSON streaming http://localhost:11434

Provider APIs remain separate. There is no provider-neutral facade that hides provider-specific capabilities or request models.

Root clients are immutable after construction and safe for concurrent use; domain clients are obtained through accessors such as client.Chat() and client.Completions().

For the current endpoint and capability status, see:

Development

Generated easyjson files are committed to the repository and must not be edited manually. Run the complete local quality workflow:

go generate ./...
make lint
make tests
go test -race ./...
go vet ./...
go mod verify
git diff --check

Tests use deterministic mocks and do not require provider credentials or live network access. See AGENTS.md for architecture, dependency, serialization, security, and resource-lifecycle rules.

Contributing

Contributions are welcome. Before opening a pull request:

  1. Keep provider-specific types and operations in their provider/domain package.
  2. Keep shared transport, auth, error, pagination, stream, and WebSocket lifecycle code in pkg without provider imports.
  3. Regenerate easyjson output with go generate ./....
  4. Run the development checks listed above.
  5. Add deterministic tests for new protocol behavior and resource ownership.

License

This project is available under the BSD 3-Clause License.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages