FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

inference-serving · GitHub Topics · GitHub

#

inference-serving

Here are 18 public repositories matching this topic...

[SIGCOMM 2023] Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient Inference

  • Updated Nov 17, 2023
  • Verilog

[Long Term Support] [SIGCOMM 2023] Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient Inference

  • Updated Sep 20, 2024
  • Verilog

The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"

  • Updated Mar 17, 2025

Official repo for the ACSOS 2021 paper on how to manage many deep learning models at the edge!

  • Updated Sep 29, 2021
  • Python

NexusKV is a model-state intelligence layer for inference systems.

  • Updated Aug 21, 2026
  • Python

Local inference serving with adaptive batching, benchmark sweeps, and regression gates for batching tradeoffs.

  • Updated May 20, 2026
  • Python

Locus is an engine-neutral inference control plane for compute and model-state placement.

  • Updated Aug 22, 2026
  • Rust

Reproducible Apple-Silicon (Apple MLX) deployment and one-token verification kit for the FreeToken edge-native MoE serving API — pinned, loopback-only, schema-validated evidence.

  • Updated Aug 29, 2026
  • Python

GPU-aware LLM serving runtime with async scheduling, dynamic batching, streaming, admission control, metrics, and mock/llama.cpp/vLLM backends.

  • Updated Jul 21, 2026
  • Python

Machine-readable companion to the IEEE OJ-CS survey 'Semantic Caching and Response Reuse for Large Language Model Services: A Survey' (Chukkapalli, Mishra, Naik, 2026): 21-work evidence matrix, systematic-search log, proposed benchmark trace schema, stdlib-only contract validator, and CPU pilot. Code MIT; data CC-BY-4.0.

  • Updated Jul 11, 2026
  • Python

GSM artifact for exact online GNN inference serving with degree-based scheduling and memory-aware batch division.

  • Updated Jun 15, 2026
  • Python

Put AI compute in the 82 million houses that already have a grid connection. An exact-arithmetic feasibility study of residential AI inference as an alternative to the data center buildout.

  • Updated Aug 24, 2026
  • HTML

Evidence-first FastAPI inference gateway for guarded LLM and RAG workloads.

  • Updated Aug 12, 2026
  • Python

cost-aware-inference-cluster: FastAPI inference-serving prototype with Redis queues, worker heartbeats, dynamic batching, autoscaling simulation, and local Docker Compose benchmark evidence.

  • Updated Jun 29, 2026
  • Python

Your GPU is not 90% busy. truthscale measures what each node can actually deliver, reports the gap, and makes that the signal you scale on.

  • Updated Aug 20, 2026
  • Go

Helm chart for deploying vLLM on GPU Kubernetes — startup probes tuned for slow model loads, correct GPU scheduling, sized /dev/shm, and GPU-utilization/queue-depth autoscaling.

  • Updated Aug 28, 2026
  • Go Template

Inference gateway for batching AI model requests through bounded micro-batching, queue backpressure, per-request result routing, and reproducible performance evaluation.

  • Updated Aug 4, 2026
  • TypeScript

Improve this page

Add a description, image, and links to the inference-serving topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the inference-serving topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL