| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuseNCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.
Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.
Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruction.
Several optimization methods of half-precision general matrix vector multiplication (HGEMV) using CUDA core.
Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.
| Back | FazBrowse Home | New Git URL |