FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Original HTTPS Page]

xlite-dev Β· GitHub

xlite-dev

Develop ML/AI toolkits and ML/AI/CUDA Learning resources.

Latest News

  • [2026-08] 🐍 Cache-DiT x FFPA (FP8/FP4) is ready! Feel free to take a try for your Diffusion models. πŸŽ‰πŸŽ‰
  • [2026-08] πŸšͺ FFPA now experimental supports FP4 Attention for headdims [64,1024] (sm_120, forward only), achieving 850-980πŸŽ‰ TFLOPS (D=128-256) on NVIDIA RTX 5090, 3.8x~4.4xπŸŽ‰ speedup over PyTorch SDPA (FlashAttention-2 backend), the performance of large headdims is stay tuned for updates. πŸŽ‰πŸŽ‰
  • [2026-08] πŸ¦… FFPA now supports D=512 for NVIDIA B200 via CuTe-DSL tcgen05 2-CTA, 1517 TFLOPS forward and 763 TFLOPS backward, achieving 6x~15xπŸŽ‰ speedup over standard PyTorch SDPA. πŸŽ‰πŸŽ‰
  • [2026-07] 🎯 FFPA now supports FP8 Attention for headdims [64,1024] (sm_120, forward only) and achieving 3x~6xπŸŽ‰ speedup over PyTorch SDPA for large headdim (D>256). πŸŽ‰πŸŽ‰
  • [2026-06] FFPA now supports AMD ROCm/HIP GPUs via the TritonBackend, check #268 for more details. πŸŽ‰
  • [2026-06] πŸ¦… NVIDIA-Nemo/AutoModel x FFPA achieving 1.4x~1.5xπŸŽ‰ End2End training throughput speedup for Gemma4-31B (8xH200, FSDP2 + AC) with FFPA accelerating the 10/60 (D=512) full-attention layers. πŸŽ‰πŸŽ‰
  • [2026-06] 🐍 FFPA now supports TritonBackend and CuTeDSLBacked for both forward and backward pass, achieving 1.5x~5xπŸŽ‰ speedup over standard PyTorch SDPA across many devices. πŸŽ‰πŸŽ‰
  • [2026-05] πŸšͺ FFPA now supports GQA, MQA, cross-attn, causal, attn-mask and dropout with CUDABackend for large headdims (D>256, forward only), achieving 1.3x~2xπŸŽ‰ speedup over PyTorch SDPA. πŸŽ‰πŸŽ‰

Pinned Loading

  1. LeetCUDA LeetCUDA Public

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    Cuda 11.8k 1.2k

  2. lite.ai.toolkit lite.ai.toolkit Public

    A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.

    C++ 4.4k 783

  3. Awesome-LLM-Inference Awesome-LLM-Inference Public

    πŸ“šA curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.πŸŽ‰

    Python 5.5k 429

  4. Awesome-DiT-Inference Awesome-DiT-Inference Public

    πŸ“šA curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.πŸŽ‰

    Python 587 29

  5. torchlm torchlm Public

    πŸ’ŽAn easy-to-use PyTorch library for face landmarks detection: training, evaluation, inference, and 100+ data augmentations.πŸŽ‰

    Python 270 29

  6. ffpa-attn ffpa-attn Public

    Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.

    Python 323 25

Repositories

Loading
Type
Select type
All Public Sources Forks Archived Mirrors Templates
Language
Select language
All C++ Cuda HTML Python Shell TypeScript
Sort
Select order
Last updated Name Stars
Showing 10 of 73 repositories

Top languages

Loading…

Most used topics

Loading…


Back | FazBrowse Home | New Git URL