FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

quantization · GitHub Topics · GitHub

#

quantization

Here are 2,612 public repositories matching this topic...

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

  • Updated Aug 27, 2026
  • Python

Faster Whisper transcription with CTranslate2

  • Updated Nov 19, 2025
  • Python

中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)

  • Updated Apr 19, 2026
  • Python

[🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment Research Platform. 📃 online docs: https://ufund-me.github.io/Qbot ✨ :news: qbot-mini: https://github.com/Charmve/iQuant

  • Updated Mar 11, 2026
  • Jupyter Notebook

A vector index built on TurboQuant, written in Rust with Python bindings

  • Updated Aug 21, 2026
  • Rust

Accessible large language models via k-bit quantization for PyTorch.

  • Updated Aug 27, 2026
  • Python

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

  • Updated Aug 26, 2026
  • C

Lossy PNG compressor — pngquant command based on libimagequant library

  • Updated Jun 21, 2026
  • C

An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.

  • Updated Apr 11, 2025
  • Python

[ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models

  • Updated Mar 7, 2026
  • Python

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

  • Updated Aug 28, 2026
  • Python

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

  • Updated Jan 17, 2026
  • Cuda

🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools

  • Updated Aug 24, 2026
  • Python

Pretrained language model and its related optimization techniques developed by Huawei Noah's Ark Lab.

  • Updated Jan 22, 2024
  • Python

PyTorch native quantization for training and inference

  • Updated Aug 28, 2026
  • Python

A model library for exploring state-of-the-art deep learning topologies and techniques for optimizing Natural Language Processing neural networks

  • Updated Nov 7, 2022
  • Python

ComfyUI Plugin of Nunchaku

  • Updated Feb 19, 2026
  • Python

Your Cheat Sheet for AI Engineering Interview – Questions and Answers.

  • Updated Aug 27, 2026
  • Markdown

Improve this page

Add a description, image, and links to the quantization topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the quantization topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL