FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Original HTTPS Page]

xlite-dev · GitHub

xlite-dev

Develop ML/AI toolkits and ML/AI/CUDA Learning resources.

ffpa-attn: Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x🔥🔥 speedup over standard PyTorch SDPA.


BF16 Attention for Large Headdim: FFPA vs SDPA (FWD/BWD) across NVIDIA H200 and B200, 6x-15x↑.


FP4 Attention for D=128: FFPA vs SageAttention-3 (FWD) on NVIDIA RTX PRO 6000.

Pinned Loading

  1. LeetCUDA LeetCUDA Public

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    Cuda 11.8k 1.2k

  2. lite.ai.toolkit lite.ai.toolkit Public

    A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.

    C++ 4.4k 785

  3. Awesome-LLM-Inference Awesome-LLM-Inference Public

    📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

    Python 5.5k 430

  4. Awesome-DiT-Inference Awesome-DiT-Inference Public

    📚A curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.🎉

    Python 589 29

  5. torchlm torchlm Public

    💎An easy-to-use PyTorch library for face landmarks detection: training, evaluation, inference, and 100+ data augmentations.🎉

    Python 270 29

  6. ffpa-attn ffpa-attn Public

    Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.

    Python 328 27

Repositories

Loading
Type
Select type
All Public Sources Forks Archived Mirrors Templates
Language
Select language
All C++ Cuda HTML Python Shell TypeScript
Sort
Select order
Last updated Name Stars
Showing 10 of 74 repositories

Top languages

Loading…

Most used topics

Loading…


Back | FazBrowse Home | New Git URL