| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
LiteLLM for Go. A single provider-agnostic client over Anthropic, OpenAI, Gemini, and the rest — with a router, fallback, cost tracking, and capability flags on top.
llmgate is the LLM gateway that vectorless-engine depends on, extracted into its own module so anything written in Go can use it. It is not a rewrite of LiteLLM. It sits on top of tmc/langchaingo — langchaingo handles every provider's wire protocol, llmgate wraps that behind a tiny Client interface and adds the production features langchaingo deliberately doesn't include.
The caller holds a Client. Middlewares (retry.New, budget.New, cache.New) wrap the client; a router.New can sit anywhere in the chain to fall over between providers. The private internal/adapter is the single seam where llmgate's interface meets langchaingo's provider implementations — one adapter serves all three providers.
Order matters. The outermost wrapper sees the call first; the innermost hits the network. Put cache.New below budget.New so cache hits don't burn budget. Put retry.New on top so retries run regardless of which inner layer tripped.
A call threads through the stack, optionally falls over to a backup provider, and comes back with Usage.{InputTokens, OutputTokens, TotalTokens, CostUSD} populated — no extra call, computed from a static price table.
The Go ecosystem has two extremes:
llmgate is the middle layer. One interface. Every provider behind it. All the production concerns — router, fallback on rate-limit, cost per call, capability flags — baked in rather than bolted on.
Early code. Interface is the stable surface; implementations evolve underneath. Roadmap lives in ROADMAP.md.
What's in now:
Client interface with Complete + CountTokens
Anthropic, OpenAI, Gemini — all backed by langchaingo/llms, in the provider/ subpackages
A single internal adapter; add a provider = add a ~30-line file
retry.New middleware for exp-backoff on transient errors
Cost tracking via a static pricing table + Usage.CostUSD on every Response
Capability flags (MaxContext, SupportsJSONMode, SupportsStreaming, SupportsTools, SupportsVision) with a capabilities.Capable interface
router.New with per-provider fallback (router.OnRateLimit, router.OnTransient, or a custom router.FallbackPolicy)
budget.New middleware with daily + total USD caps and UTC rollover
cache.New middleware — in-memory LRU keyed on request shape, optional TTL
Error classification (Classify, IsRateLimited, IsTransient, IsAuth) that the retry predicate and router policies use to decide what to retry or fall over on
Mock client with call recording for tests
Native tool calling (Request.Tools, Response.ToolCalls) across all three providers
Coming next:
go get github.com/hallelx2/llmgatepackage main
import (
"context"
"fmt"
"os"
"github.com/hallelx2/llmgate"
"github.com/hallelx2/llmgate/middleware/retry"
"github.com/hallelx2/llmgate/provider/anthropic"
)
func main() {
client, err := anthropic.New(anthropic.Config{
APIKey: os.Getenv("ANTHROPIC_API_KEY"),
Model: "claude-sonnet-4-5",
})
if err != nil {
panic(err)
}
// Optional: wrap in exponential-backoff retry middleware.
client = retry.New(retry.Config{MaxRetries: 3})(client)
resp, err := client.Complete(context.Background(), llmgate.Request{
Messages: []llmgate.Message{
{Role: llmgate.RoleUser, Content: "In one sentence: what is vectorless retrieval?"},
},
MaxTokens: 256,
})
if err != nil {
panic(err)
}
fmt.Println(resp.Content)
}Every Response carries a Usage with the token breakdown split by billing tier, not just a prompt/completion pair:
resp.Usage.InputTokens // uncached prompt tokens
resp.Usage.CacheWriteTokens // written to the prompt cache (1.25x on Anthropic)
resp.Usage.CacheReadTokens // served from cache (0.1x Anthropic, 0.5x OpenAI, 0.25x Google)
resp.Usage.ReasoningTokens // thinking tokens — a subset of OutputTokens
resp.Usage.CostUSDThe providers disagree about whether cached tokens are counted inside the prompt total — Anthropic reports them alongside, OpenAI and Google report them within — so llmgate normalizes to the disjoint form. InputTokens + CacheWriteTokens + CacheReadTokens is the whole prompt on every provider.
Three flags say how much to trust the number:
| Flag | Meaning when false |
|---|---|
| Priced | no price-book entry — CostUSD is unknown, not zero |
| TokensReported | the provider returned no counts |
| Estimated | (when true) counts came from a local tokenizer, so cost is approximate |
That distinction matters: a CostUSD of 0 with Priced true asserts the call was free, which it never is. When a provider reports no usage at all, llmgate estimates from a tokenizer and flags it rather than reporting zero.
Lookups resolve dated, prefixed, and gateway-qualified IDs to their base model, so none of these price at $0:
claude-sonnet-4-5-20250929 -> claude-sonnet-4-5 models/gemini-2.5-flash -> gemini-2.5-flash us.anthropic.claude-opus-4-1-v1:0 -> claude-opus-4-1 z-ai/glm-4.6 -> glm-4.6
Rates are keyed by model ID alone, never by provider — a model served through a gateway that speaks another vendor's protocol (GLM over an Anthropic-compatible endpoint, say) still prices correctly.
The compiled-in table drifts as vendors change rates. No vendor publishes machine-readable pricing, so UseRemote layers a community aggregate over it:
stop, err := pricing.UseRemote(ctx, pricing.RemoteConfig{
CacheDir: "/var/cache/llmgate", // survive restarts
OnError: func(src string, err error) { log.Warn("price refresh", "src", src, "err", err) },
})
defer stop()Sources default to LiteLLM's price table then OpenRouter's model API, both public and unauthenticated.
This is opt-in and stays that way — importing pricing performs no network I/O. Resolution is Register overrides → remote snapshot → embedded table, and it fails open at every layer: lookups never block on the network, a failed fetch keeps the previous snapshot, and a snapshot whose rates have drifted more than 10x from the embedded values is rejected as corrupt rather than adopted. pricing.AsOf() reports the vintage of whatever is loaded.
See ROADMAP.md for what is shipped and what is next.
Apache 2.0. See LICENSE.
| Back | FazBrowse Home | New Git URL |