FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

cuda-optimization · GitHub Topics · GitHub

#

cuda-optimization

Here are 7 public repositories matching this topic...

CUDA matrix multiplication benchmarking on Jetson Orin Nano. Four implementations, three power modes, five matrix sizes. 99.5% mathematical validation. C++/CUDA and Python.

  • Updated Apr 2, 2026
  • Python

FlashAttention-style CUDA implementation with shared-memory tiling, online softmax fusion, IO-aware optimization, and GPU benchmarking.

  • Updated Aug 13, 2026
  • Python

DeepSeek-R1 7B INT4 at 69.3 tok/s on a $300 RTX 3060. Faster than llama.cpp, vLLM, and NVIDIA TensorRT-LLM. Is one developer + Ai really better than the entire industry?

  • Updated May 19, 2026
  • Python

A 110M-parameter Llama-style transformer trained from scratch on the TinyStories dataset, optimized for high-throughput training on 4GB VRAM consumer GPUs. The project features a custom asynchronous CUDA-stream prefetcher and KV-cache inference, achieving 10k+ TPS on an RTX 3050.

  • Updated Apr 8, 2026
  • Python

🏆 Which 3DGS renderer is fastest? Which compression is best? We measured them all — on the same GPU, same scenes, same protocol.

  • Updated Aug 23, 2026
  • Python

Hardened RAG pipeline with Llama 3.2 (3B) & Arize Phoenix. Features 4-bit Unsloth optimization, OpenTelemetry auditing, and a KV-cache stability patch for T4 GPUs. P99 Latency: 19.2s.

  • Updated Mar 31, 2026
  • Python

Improve this page

Add a description, image, and links to the cuda-optimization topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the cuda-optimization topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL