| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Open-source AI gateway for LLMs & AI agents, built in Rust. One OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock & more — routing, guardrails, caching, rate limits, observability.
An Agent Development Kit (ADK) allowing for seamless creation of A2A-compatible agents written in Go.
A2A agent server enabling Google Calendar scheduling, retrieval, and automation
An Agent Development Kit (ADK) allowing for seamless creation of A2A-compatible agents written in Rust.
A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
OpenAI-compatible inference coordinator for local llama.cpp workers — multi-worker scheduling, systemd lifecycle, streaming, tool calls.
A Git-first CLI coding agent that turns ideas, issues, and tasks into real code changes. It can run remotely from your mobile phone or automate your computer.
An SDK written in Go for the Inference Gateway.
ROCm-first OpenAI-compatible inference gateway for open-weight reasoning models. Single active profile (gpt-oss-20b, deepseek-r1-distill, qwen3-4b), host-mounted weights, fail-closed contract. FastAPI + vLLM.
An intelligent gateway for Claude APIs that dynamically routes requests to the most cost-efficient model, caches responses, and escalates based on confidence signals — reducing LLM spend without sacrificing quality.
Inference Ops exercise for my class
🔌 单卡 GPU LLM 推理网关 · 模型即插件 · 三态 GPU · 9 云端预设 · macOS Dashboard
An SDK written in Rust for the Inference Gateway
Deterministic local-first inference gateway that controls cost, caching, and routing for predictable AI execution.
Curated catalog of Agent Skills for the Inference Gateway ecosystem
An enterprise-grade, configuration-driven MLOps pipeline for credit risk underwriting. Built with XGBoost, strict data validation, mlFlow, and CI/CD automation. Dockerized inference deployed via render
Extensive documentation of the inference-gateway.
Production-grade Java 25 Virtual Thread inference gateway bridging NVIDIA Triton → Dynamo with Earliest Deadline First (EDF) priority queuing, adaptive batching, and async shadow validation.
The blood-brain barrier for autonomous agents. A context-aware LLM governance proxy that enforces credential starvation — identity-verified, provider-routed, cost-tracked, and audit-logged.
Add a description, image, and links to the inference-gateway topic page so that developers can more easily learn about it.
To associate your repository with the inference-gateway topic, visit your repo's landing page and select "manage topics."
| Back | FazBrowse Home | New Git URL |