FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

tgi · GitHub Topics · GitHub

#

tgi

Here are 16 public repositories matching this topic...

Generative AI Examples is a collection of GenAI examples such as ChatQnA, Copilot, which illustrate the pipeline capabilities of the Open Platform for Enterprise AI (OPEA) project.

  • Updated Aug 20, 2026
  • Shell

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

  • Updated Aug 28, 2026
  • Go

大模型推理框架加速,让 LLM 飞起来

  • Updated May 10, 2024
  • Python

Bench360 is a modular benchmarking suite for local LLM deployments. It offers a full-stack, extensible pipeline to evaluate the latency, throughput, quality, and cost of LLM inference on consumer and enterprise GPUs. Bench360 supports flexible backends, tasks and scenarios, enabling fair and reproducible comparisons for researchers & practitioners.

  • Updated Feb 18, 2026
  • Python

The control plane for self-hosted AI inference. Warm-state GPU routing, multi-runtime orchestration across Ollama, vLLM, llama.cpp, TGI and MLX . Single Go binary. Apache-2.0.

  • Updated Aug 28, 2026
  • Go

AWS deployment stack for Gemma 3 on SageMaker with HuggingFace TGI, OpenAI-compatible API (Lambda + API Gateway), and OpenWebUI chat interface

  • Updated Mar 25, 2026
  • Python

Zitate & Memes der Klasse TGI13

  • Updated May 17, 2019

LLM Inference performance harness

  • Updated Dec 29, 2025
  • Python

Bridge GitHub Copilot Chat with local vLLM/TGI servers and HuggingFace cloud models. Enterprise-ready VS Code extension for air-gapped AI coding.

  • Updated Sep 29, 2025
  • TypeScript

An nvtop for local LLM inference: zero-config autodiscovery of vLLM, llama.cpp, Ollama, TGI, SGLang + live GPU and serving metrics in a Textual TUI. Unified-memory (GB10/Jetson) aware. MIT.

  • Updated Jul 7, 2026
  • Python

TGI server setup for Intel Data Centre GPUs

  • Updated Nov 26, 2024
  • Shell

Self-hosted FastAPI gateway exposing OpenAI and Anthropic Messages APIs in front of any open-source LLM runtime (vLLM, Ollama, llama.cpp, TGI, SGLang, LocalAI, LM Studio). Streaming, embeddings, metrics, auth, rate limiting.

  • Updated Apr 22, 2026
  • Python

Comprehensive LLM deployment analysis: vLLM, TGI, Ollama, and production-ready strategies

  • Updated Jul 12, 2026

Nemotron-3-Ultra-550B, Gemma4-31B on 8xA100 40GB GPUs as production AWS SageMaker endpoint with tensor-parallel inference, Gemma4-12B-it on Lambda GPU, Megatron-Bridge w vLLM or SGLang,

  • Updated Jul 29, 2026
  • Python

it's the .github repo 🚀

tgi
  • Updated Sep 30, 2025

Lightweight HTML form with Python Flask app and accompanying scripts for swift testing of interactions with SEA-LION family of LLMs.

  • Updated Aug 2, 2024
  • Python

Improve this page

Add a description, image, and links to the tgi topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the tgi topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL