| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming
Implementation of FusedMM method for IPDPS 2021 paper titled "FusedMM: A Unified SDDMM-SpMM Kernel for Graph Embedding and Graph Neural Networks"
Custom CUDA & Triton fused layers for self-stabilizing transformer architectures. Accelerate forward/backward passes and prevent gradient explosions in large-scale LLM training.
Contents for my Masters thesis "Optimized Block-Level Matrix Inversion Kernels for Small, Batched Matrices on GPUs"
Add a description, image, and links to the fused-kernel topic page so that developers can more easily learn about it.
To associate your repository with the fused-kernel topic, visit your repo's landing page and select "manage topics."
| Back | FazBrowse Home | New Git URL |