| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
NumPy & SciPy for GPU
Hooked CUDA-related dynamic libraries by using automated code generation tools.
A safe Rust FFI binding for the NVIDIA® Tools Extension SDK (NVTX).
PyProf2: PyTorch Profiling tool
Julia bindings for NVTX, for instrumenting with the Nvidia Nsight Systems profiler
Simple timing routines to be used in codes which use MPI and possibly CUDA/OpenACC using NVTX markers
Thin pybind11 wrapper for NVTX wrappers -- with some bells and whistles attached.
Create NVIDIA NVTX ranges directly in VS Code, then profile with Nsight Systems without modifying source code.
Reproducible GPT-2 distributed-training benchmarks on 1-8 V100 GPUs using Slurm, PyTorch, DeepSpeed, NCCL, NVTX, and Nsight Systems.
🎬 Explore GPU training efficiency with FP32 vs FP16 in this modular lab, utilizing Tensor Core acceleration for deep learning insights.
A reproducible GPU benchmarking lab that compares FP16 vs FP32 training on MNIST using PyTorch, CuPy, and Nsight profiling tools. This project blends performance engineering with cinematic storytelling—featuring NVTX-tagged training loops, fused CuPy kernels, and a profiler-driven README that narrates the GPU’s inner workings frame by frame.
Profile-first ML systems project optimizing a multi-camera end-to-end driving model for hardware efficiency using PyTorch, CUDA streams, NVTX instrumentation, and Nsight Systems.
Header-only C++17 latency profiler for real-time CUDA loops. Per-stage distributions, deadline tracking, and Perfetto timelines
Profiling with Precision. Documenting with Style.
Add a description, image, and links to the nvtx topic page so that developers can more easily learn about it.
To associate your repository with the nvtx topic, visit your repo's landing page and select "manage topics."
| Back | FazBrowse Home | New Git URL |