FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

gpu-benchmarking · GitHub Topics · GitHub

#

gpu-benchmarking

Here are 17 public repositories matching this topic...

CUDA matrix multiplication benchmarking on Jetson Orin Nano. Four implementations, three power modes, five matrix sizes. 99.5% mathematical validation. C++/CUDA and Python.

  • Updated Apr 2, 2026
  • Python

Dashboard for AI Studio, Open Source Continuous Inference | Deepseek-R1, Qwen2.5, Llama3.1 | 4xRTX-5090 inside PRU2500, 2xH100 inside PRU2500, 8xMI210 in SuperMicro

  • Updated Aug 7, 2026
  • TypeScript

LLM benchmarking, GPU workload orchestration backend server | Deepseek-R1, Qwen2.5, Llama3.1 | 4xRTX-5090 inside PRU2500, 2xH100 inside PRU2500, 8xMI210 in SuperMicro

  • Updated Aug 18, 2026
  • Python

Automated game testing and GPU Cooling Tester Evaluation with reproducible measurements and profile authoring

  • Updated Aug 27, 2026
  • C#

Standalone LLM inference benchmarking pipelines on AMD GPUs using ROCm, vLLM, MAD, and data visualization scripts.

  • Updated Feb 21, 2026
  • Python

Hands-on Jupyter notebooks for deep learning with TensorFlow, covering fundamental concepts, model training, and applied tabular projects.

  • Updated May 29, 2026
  • Jupyter Notebook

Plots GPU benchmarks across user-selected axes and highlights the Pareto frontier to help builders choose hardware for local AI inference without editorial bias.

  • Updated Apr 18, 2026
  • HTML

Artifact-backed LLM serving performance lab for vLLM baselines, official metrics, GuideLLM checks, and SGLang/PD scaffolding

  • Updated May 21, 2026
  • Python
  • Updated Jan 19, 2022
  • C++

benchHUB is a Python-based project to parse, aggregate, and visualize system and performance benchmarks. It includes a Streamlit dashboard to display and compare results.

  • Updated Jun 19, 2026
  • Python

Reproducible GPT-2 distributed-training benchmarks on 1-8 V100 GPUs using Slurm, PyTorch, DeepSpeed, NCCL, NVTX, and Nsight Systems.

  • Updated Jun 25, 2026
  • Python

Measured CUDA graphs trade-off in vLLM 0.23 on H100: per-token decode always faster, but throughput inverts at 32B under saturation (-5.15%) alongside -1.8% tok/J; cold-start cost decomposed into capture and torch.compile.

  • Updated Jul 26, 2026
  • Python

One-shot script to audit GPU, CUDA, PyTorch, CPU, and disk performance before debugging a slow or broken ML environment.

  • Updated Apr 3, 2026
  • Shell

Professional benchmarking guide for evaluating NVIDIA H200 GPU memory bandwidth using the STREAM Triad (FP32) benchmark with CUDA 12.x on Linux. Includes setup, execution, performance analysis, and multi-GPU benchmarking results.

  • Updated Jul 25, 2026

Run a 2-min local benchmark → predict how long your AI job will take on cloud GPU (T4/V100/A100). No guessing, no wasted money.

  • Updated Aug 17, 2026
  • Python

Ramanujan's mathematics meets the NVIDIA stack: CUDA-Q/cuQuantum quantum simulation + NIM/Nemotron analysis, consumer RTX to cloud H100

  • Updated Aug 5, 2026
  • Python

Evidence-backed, workload-conditioned autotuning for local LLM inference. Local-first control plane, G0–G8 validation, PyTorch · llama.cpp · vLLM.

  • Updated Aug 11, 2026
  • Python

Improve this page

Add a description, image, and links to the gpu-benchmarking topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the gpu-benchmarking topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL