FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

inference-efficiency · GitHub Topics · GitHub

#

inference-efficiency

Here are 11 public repositories matching this topic...

Official code for paper: [CLS] Attention is All You Need for Training-Free Visual Token Pruning: Make VLM Inference Faster.

  • Updated Jun 29, 2025
  • Python

[NeurIPS 2025] Official code for paper: Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

  • Updated Sep 20, 2025
  • Python

[TCSVT] Official repository of the paper "A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models"

  • Updated Jun 12, 2026
  • Python

This is a collection of our research on efficient AI, covering hardware-aware NAS and model compression.

  • Updated Oct 25, 2024
  • Python

[ICCV 2025] Official code for paper: Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

  • Updated Jul 1, 2025
  • Python

A tutorial of model quantization using TensorFlow

  • Updated Aug 2, 2021
  • Python

DWARF v2 is a research prototype that expands on the original DWARF.

  • Updated Aug 24, 2026
  • Python

O(N) attention with a bounded inference KV cache. D4 Daubechies wavelet field + content-gated Q·K gather at dyadic offsets.

  • Updated Jun 9, 2026
  • Python

Substrate-agnostic, structure-complete memory fabric — read the whole book from any address; forward-compatible across classical, photonic, and quantum hardware (rare-earth quantum memory). Reproducible O(N^2)->O(N) LLM efficiency.

  • Updated Jul 9, 2026
  • Python

Reproducibility materials for causal whitespace patching in Korean byte-latent LMs: a 2.5–2.9% matched-quality latency result without scale amplification.

  • Updated Aug 17, 2026
  • Python

Official implementation of recall-controlled early abort for LLM agent episodes

  • Updated Aug 18, 2026
  • Python

Improve this page

Add a description, image, and links to the inference-efficiency topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the inference-efficiency topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL