FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

multi-gpu-inference · GitHub Topics · GitHub

#

multi-gpu-inference

Here are 7 public repositories matching this topic...

Language: All
Filter by language

Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).

  • Updated Mar 15, 2024
  • C++

A script for PyTorch multi-GPU multi-process testing

  • Updated Apr 29, 2024
  • Python

Distributed Reinforcement Learning for LLM Fine-Tuning with multi-GPU utilization

  • Updated Mar 12, 2025
  • Python

MiniMax-H3 multi-GPU parallel inference for ComfyUI | 多卡并行加速节点:2-8 GPU Ulysses sequence parallel, bit-identical video+audio generation

  • Updated Aug 14, 2026
  • Python

A reproducible Docker-based pipeline for running machine learning experiments with GPU support. This repository provides pre-configured Docker images, environment files, and scripts for: - Setting up GPU-enabled containers - Running training and inference - Managing environments reproducibly

  • Updated Nov 14, 2025
  • Dockerfile

Nemotron-3-Ultra-550B, Gemma4-31B on 8xA100 40GB GPUs as production AWS SageMaker endpoint with tensor-parallel inference, Gemma4-12B-it on Lambda GPU, Megatron-Bridge w vLLM or SGLang,

  • Updated Jul 29, 2026
  • Python

Lightweight terminal launcher and auto-optimizer for llm models using llama.cpp with hardware detection, tensor sharding, benchmarks, presets, context tuning, and OpenAI-compatible serving.

  • Updated Aug 11, 2026
  • Python

Improve this page

Add a description, image, and links to the multi-gpu-inference topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the multi-gpu-inference topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL