| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Reference implementations of MLPerf® inference benchmarks
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.
The plugin-driven server agent for collecting & reporting metrics.
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |