FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

sglang · GitHub Topics · GitHub

#

sglang

Here are 249 public repositories matching this topic...

A Datacenter Scale Distributed Inference Serving Framework

  • Updated Aug 20, 2026
  • Rust

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

  • Updated Aug 19, 2026
  • C++

OpenClaw-RL: Train any agent simply by talking

  • Updated May 23, 2026
  • Python

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

  • Updated Aug 19, 2026
  • Python

Control panel for VLLM, Sglang, llama.cpp, exllamav3

  • Updated Aug 19, 2026
  • TypeScript

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

  • Updated Aug 19, 2026
  • Python

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

  • Updated Aug 19, 2026
  • Python

MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.

  • Updated Jul 26, 2026
  • Python

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

  • Updated Aug 19, 2026
  • Python

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

  • Updated Aug 19, 2026
  • Python

MOVA: Towards Scalable and Synchronized Video–Audio Generation

  • Updated Aug 19, 2026
  • Python

UniRL is a Framework for Unified Multimodal Model Reinforcement Learning

  • Updated Aug 19, 2026
  • Python

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

  • Updated Aug 19, 2026
  • Python

A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

  • Updated Aug 14, 2026
  • HTML

基于SparkTTS、OrpheusTTS等模型,提供高质量中文语音合成与声音克隆服务。

  • Updated May 18, 2025
  • Python

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

  • Updated Aug 19, 2026
  • Go

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

  • Updated Aug 20, 2026
  • Rust

sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems

  • Updated Aug 18, 2026
  • Python

☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!

  • Updated Jan 26, 2026
  • Go

Improve this page

Add a description, image, and links to the sglang topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the sglang topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL