FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

inference-gateway · GitHub Topics · GitHub

#

inference-gateway

Here are 32 public repositories matching this topic...

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

  • Updated Aug 28, 2026
  • Rust

Open-source AI gateway for LLMs & AI agents, built in Rust. One OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock & more — routing, guardrails, caching, rate limits, observability.

  • Updated Aug 28, 2026
  • Rust

An Agent Development Kit (ADK) allowing for seamless creation of A2A-compatible agents written in Go.

  • Updated Aug 27, 2026
  • Go

A2A agent server enabling Google Calendar scheduling, retrieval, and automation

  • Updated Aug 27, 2026
  • Go

An Agent Development Kit (ADK) allowing for seamless creation of A2A-compatible agents written in Rust.

  • Updated Aug 27, 2026
  • Rust

A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

  • Updated Aug 27, 2026
  • Go

OpenAI-compatible inference coordinator for local llama.cpp workers — multi-worker scheduling, systemd lifecycle, streaming, tool calls.

  • Updated Aug 2, 2026
  • Rust

A Git-first CLI coding agent that turns ideas, issues, and tasks into real code changes. It can run remotely from your mobile phone or automate your computer.

  • Updated Aug 28, 2026
  • Go

An SDK written in Go for the Inference Gateway.

  • Updated Aug 27, 2026
  • Go

ROCm-first OpenAI-compatible inference gateway for open-weight reasoning models. Single active profile (gpt-oss-20b, deepseek-r1-distill, qwen3-4b), host-mounted weights, fail-closed contract. FastAPI + vLLM.

  • Updated Apr 26, 2026
  • Python

An intelligent gateway for Claude APIs that dynamically routes requests to the most cost-efficient model, caches responses, and escalates based on confidence signals — reducing LLM spend without sacrificing quality.

  • Updated May 6, 2026
  • Python

Inference Ops exercise for my class

  • Updated Mar 29, 2026
  • Python

🔌 单卡 GPU LLM 推理网关 · 模型即插件 · 三态 GPU · 9 云端预设 · macOS Dashboard

  • Updated Aug 26, 2026
  • Python

An SDK written in Rust for the Inference Gateway

  • Updated Aug 27, 2026
  • Rust

Deterministic local-first inference gateway that controls cost, caching, and routing for predictable AI execution.

  • Updated Apr 11, 2026
  • C#

Curated catalog of Agent Skills for the Inference Gateway ecosystem

  • Updated Aug 27, 2026
  • JavaScript

An enterprise-grade, configuration-driven MLOps pipeline for credit risk underwriting. Built with XGBoost, strict data validation, mlFlow, and CI/CD automation. Dockerized inference deployed via render

  • Updated Jun 24, 2026
  • Python

Extensive documentation of the inference-gateway.

  • Updated Aug 28, 2026
  • JavaScript

Production-grade Java 25 Virtual Thread inference gateway bridging NVIDIA Triton → Dynamo with Earliest Deadline First (EDF) priority queuing, adaptive batching, and async shadow validation.

  • Updated Jun 13, 2026
  • Java

The blood-brain barrier for autonomous agents. A context-aware LLM governance proxy that enforces credential starvation — identity-verified, provider-routed, cost-tracked, and audit-logged.

  • Updated Aug 25, 2026
  • Go

Improve this page

Add a description, image, and links to the inference-gateway topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the inference-gateway topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL