| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Generative AI Examples is a collection of GenAI examples such as ChatQnA, Copilot, which illustrate the pipeline capabilities of the Open Platform for Enterprise AI (OPEA) project.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
大模型推理框架加速,让 LLM 飞起来
Bench360 is a modular benchmarking suite for local LLM deployments. It offers a full-stack, extensible pipeline to evaluate the latency, throughput, quality, and cost of LLM inference on consumer and enterprise GPUs. Bench360 supports flexible backends, tasks and scenarios, enabling fair and reproducible comparisons for researchers & practitioners.
The control plane for self-hosted AI inference. Warm-state GPU routing, multi-runtime orchestration across Ollama, vLLM, llama.cpp, TGI and MLX . Single Go binary. Apache-2.0.
AWS deployment stack for Gemma 3 on SageMaker with HuggingFace TGI, OpenAI-compatible API (Lambda + API Gateway), and OpenWebUI chat interface
LLM Inference performance harness
Bridge GitHub Copilot Chat with local vLLM/TGI servers and HuggingFace cloud models. Enterprise-ready VS Code extension for air-gapped AI coding.
An nvtop for local LLM inference: zero-config autodiscovery of vLLM, llama.cpp, Ollama, TGI, SGLang + live GPU and serving metrics in a Textual TUI. Unified-memory (GB10/Jetson) aware. MIT.
TGI server setup for Intel Data Centre GPUs
Self-hosted FastAPI gateway exposing OpenAI and Anthropic Messages APIs in front of any open-source LLM runtime (vLLM, Ollama, llama.cpp, TGI, SGLang, LocalAI, LM Studio). Streaming, embeddings, metrics, auth, rate limiting.
Comprehensive LLM deployment analysis: vLLM, TGI, Ollama, and production-ready strategies
Nemotron-3-Ultra-550B, Gemma4-31B on 8xA100 40GB GPUs as production AWS SageMaker endpoint with tensor-parallel inference, Gemma4-12B-it on Lambda GPU, Megatron-Bridge w vLLM or SGLang,
Lightweight HTML form with Python Flask app and accompanying scripts for swift testing of interactions with SEA-LION family of LLMs.
Add a description, image, and links to the tgi topic page so that developers can more easily learn about it.
To associate your repository with the tgi topic, visit your repo's landing page and select "manage topics."
| Back | FazBrowse Home | New Git URL |