FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

tensorrt-llm · GitHub Topics · GitHub

#

tensorrt-llm

Here are 66 public repositories matching this topic...

A Datacenter Scale Distributed Inference Serving Framework

  • Updated Aug 20, 2026
  • Rust

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

  • Updated Aug 14, 2026
  • Python

A nearly-live implementation of OpenAI's Whisper.

  • Updated Aug 4, 2026
  • Python

An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine

  • Updated Aug 27, 2024
  • Jupyter Notebook

🚀🚀🚀 This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PTX and High Performance Computing (HPC) projects.

  • Updated Aug 2, 2025

🏋️ A unified multi-backend utility for benchmarking Transformers, Timm, PEFT, Diffusers and Sentence-Transformers with full support of Optimum's hardware optimizations & quantization schemes.

  • Updated May 26, 2026
  • Python

OpenAI compatible API for TensorRT LLM triton backend

  • Updated Aug 1, 2024
  • Rust

Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.

  • Updated May 8, 2025
  • C++

Higher performance OpenAI LLM service than vLLM serve: A pure C++ high-performance OpenAI LLM service implemented with GPRS+TensorRT-LLM+Tokenizers.cpp, supporting chat and function call, AI agents, distributed multi-GPU inference, multimodal capabilities, and a Gradio chat interface.

  • Updated Dec 8, 2025
  • Python

TensorRT-LLM server with Structured Outputs (JSON) built with Rust

  • Updated Apr 25, 2025
  • Rust

Chat With RTX Python API

  • Updated May 11, 2025
  • Python

LLM-Inference-Bench

  • Updated Jul 18, 2025
  • Jupyter Notebook

A tool for benchmarking LLMs on Modal

  • Updated Aug 29, 2025
  • Python

Add-in for new Outlook that adds LLM new features (Composition, Summarizing, Q&A). It uses a local LLM via Nvidia TensorRT-LLM

  • Updated May 12, 2026
  • Python

Cortex.Tensorrt-LLM is a C++ inference library that can be loaded by any server at runtime. It submodules NVIDIA’s TensorRT-LLM for GPU accelerated inference on NVIDIA's GPUs.

  • Updated Sep 26, 2024
  • C++

大模型推理框架加速,让 LLM 飞起来

  • Updated May 10, 2024
  • Python

AI Infra LLM infer/ tensorrt-llm/ vllm

  • Updated Aug 10, 2026
  • Python

GPU 性能与 AI Infra 学习项目:CUDA/Triton 算子、NCU/NSYS、vLLM/SGLang/TRT-LLM/ms-swift、PyTorch/DeepSpeed/ms-swift 训练、并行架构

  • Updated Aug 18, 2026
  • Python

LLM tutorial materials include but not limited to NVIDIA NeMo, TensorRT-LLM, Triton Inference Server, and NeMo Guardrails.

  • Updated Jun 26, 2025
  • Python

Whisper in TensorRT-LLM

  • Updated Sep 21, 2023
  • C++

Improve this page

Add a description, image, and links to the tensorrt-llm topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the tensorrt-llm topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL