| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuse
I built
Cache-DiT,
ffpa-attn,
LeetCUDA,
lite.ai.toolkit,
xlite-dev, ...
🤗 I contributed to
FastDeploy,
SGLang ,
vLLM ,
Diffusers , ...
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
SGLang is a high-performance serving framework for large language models and multimodal models.
A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.
| Back | FazBrowse Home | New Git URL |