| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuse[HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.
[ACL 2024] A novel QAT with Self-Distillation framework to enhance ultra low-bit LLMs.
List of papers related to Vision Transformers quantization and hardware acceleration in recent AI conferences and journals.
CUDA Templates and Python DSLs for High-Performance Linear Algebra
| Back | FazBrowse Home | New Git URL |