FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

kvcache-optimization · GitHub Topics · GitHub

#

kvcache-optimization

Here are 8 public repositories matching this topic...

Language: All
Filter by language

vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!

  • Updated Aug 20, 2026
  • Python

KV Cache with PagedAttention vs PagedAttention + TurboQuant - experiments across token sizes comparing memory, latency, and accuracy.

  • Updated Mar 26, 2026
  • Python

[MLSys-26] FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management

  • Updated Mar 9, 2026
  • Python

Clean from-scratch inference engine for shannon-prime-lattice. NTT-based attention, two-node CRT-sharded inference path, KSTE-encoded KV state.

  • Updated Jul 10, 2026
  • HTML

Clean from-scratch math core for shannon-prime-lattice: KSTE encoder, Friedman sieve, ARM (HRR in CRT cyclotomic ring), CRT NTT primitives, Position-as-Arithmetic.

  • Updated Jul 10, 2026
  • C

Umbrella for the decentralized cooperative AI training/inference architecture built on the prime-factored coordinate lattice and the dominance order. Theory + Systems + Roadmap papers, contracts, offload pattern.

  • Updated Jul 10, 2026
  • HTML

Improve this page

Add a description, image, and links to the kvcache-optimization topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the kvcache-optimization topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL