FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

llm-serving · GitHub Topics · GitHub

#

llm-serving

Here are 366 public repositories matching this topic...

A high-throughput and memory-efficient inference and serving engine for LLMs

  • Updated Aug 28, 2026
  • Python

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

  • Updated Aug 28, 2026
  • Python

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

  • Updated Jul 19, 2026
  • HTML

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

  • Updated Aug 28, 2026
  • Python

Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

  • Updated Aug 24, 2026
  • Python

The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

  • Updated Aug 28, 2026
  • Python

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

  • Updated Aug 28, 2026
  • Python

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

  • Updated Aug 28, 2026
  • Python

Superduper: End-to-end framework for building custom AI applications and agents.

  • Updated Sep 1, 2025
  • Python

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

  • Updated May 28, 2026
  • Python

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

  • Updated Aug 26, 2026
  • Python

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

  • Updated Aug 28, 2026
  • Python

Community maintained hardware plugin for vLLM on Ascend

  • Updated Aug 28, 2026
  • C++

MoBA: Mixture of Block Attention for Long-Context LLMs

  • Updated Apr 3, 2025
  • Python

Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere

  • Updated Jul 1, 2026
  • Python

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

  • Updated Aug 28, 2026
  • Cuda

RayLLM - LLMs on Ray (Archived). Read README for more info.

  • Updated Mar 13, 2025

A throughput-oriented high-performance serving framework for LLMs

  • Updated Mar 29, 2026
  • Jupyter Notebook

Improve this page

Add a description, image, and links to the llm-serving topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the llm-serving topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL