| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
The NVIDIA Data Loading Library (DALI) is a GPU-accelerated library for data loading and pre-processing to accelerate deep learning applications. It provides a collection of highly optimized building blocks for loading and processing image, video and audio data. It can be used as a portable drop-in replacement for built in data loaders and data iterators in popular deep learning frameworks.
Deep learning applications require complex, multi-stage data processing pipelines that include loading, decoding, cropping, resizing, and many other augmentations. These data processing pipelines, which are currently executed on the CPU, have become a bottleneck, limiting the performance and scalability of training and inference.
DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the GPU. Additionally, DALI relies on its own execution engine, built to maximize the throughput of the input pipeline. Features such as prefetching, parallel execution, and batch processing are handled transparently for the user.
In addition, the deep learning frameworks have multiple data pre-processing implementations, resulting in challenges such as portability of training and inference workflows, and code maintainability. Data processing pipelines implemented using DALI are portable because they can easily be retargeted to TensorFlow, PyTorch, and PaddlePaddle.
Tip
The dali-dynamic-mode skill provides AI agents with guidance on the Dynamic Mode API and best practices. It can be installed as follows:
npx skills add nvidia/skills --skill dali-dynamic-modeFor more information, see the NVIDIA/skills GitHub repository.
Pipeline mode:
from nvidia.dali.pipeline import pipeline_def
import nvidia.dali.types as types
import nvidia.dali.fn as fn
from nvidia.dali.plugin.pytorch import DALIGenericIterator
import os
# To run with different data, see documentation of nvidia.dali.fn.readers.file
# points to https://github.com/NVIDIA/DALI_extra
data_root_dir = os.environ['DALI_EXTRA_PATH']
images_dir = os.path.join(data_root_dir, 'db', 'single', 'jpeg')
def loss_func(pred, y):
pass
def model(x):
pass
def backward(loss, model):
pass
@pipeline_def(num_threads=4, device_id=0)
def get_dali_pipeline():
images, labels = fn.readers.file(
file_root=images_dir, random_shuffle=True, name="Reader")
# decode data on the GPU
images = fn.decoders.image_random_crop(
images, device="mixed", output_type=types.RGB)
# the rest of processing happens on the GPU as well
images = fn.resize(images, resize_x=256, resize_y=256)
images = fn.crop_mirror_normalize(
images,
crop_h=224,
crop_w=224,
mean=[0.485 * 255, 0.456 * 255, 0.406 * 255],
std=[0.229 * 255, 0.224 * 255, 0.225 * 255],
mirror=fn.random.coin_flip())
return images, labels
train_data = DALIGenericIterator(
[get_dali_pipeline(batch_size=16)],
['data', 'label'],
reader_name='Reader'
)
for i, data in enumerate(train_data):
x, y = data[0]['data'], data[0]['label']
pred = model(x)
loss = loss_func(pred, y)
backward(loss, model)Dynamic mode:
import os
import nvidia.dali.types as types
import nvidia.dali.experimental.dynamic as ndd
import torch
# To run with different data, see documentation of ndd.readers.File
# points to https://github.com/NVIDIA/DALI_extra
data_root_dir = os.environ['DALI_EXTRA_PATH']
images_dir = os.path.join(data_root_dir, 'db', 'single', 'jpeg')
def loss_func(pred, y):
pass
def model(x):
pass
def backward(loss, model):
pass
reader = ndd.readers.File(file_root=images_dir, random_shuffle=True)
for images, labels in reader.next_epoch(batch_size=16):
images = ndd.decoders.image_random_crop(images, device="gpu", output_type=types.RGB)
# the rest of processing happens on the GPU as well
images = ndd.resize(images, resize_x=256, resize_y=256)
images = ndd.crop_mirror_normalize(
images,
crop_h=224,
crop_w=224,
mean=[0.485 * 255, 0.456 * 255, 0.406 * 255],
std=[0.229 * 255, 0.224 * 255, 0.225 * 255],
mirror=ndd.random.coin_flip(),
)
x = torch.as_tensor(images)
y = torch.as_tensor(labels.gpu())
pred = model(x)
loss = loss_func(pred, y)
backward(loss, model)The following issue represents a high-level overview of our 2024 plan. You should be aware that this roadmap may change at any time and the order of its items does not reflect any type of priority.
We strongly encourage you to comment on our roadmap and provide us feedback on the mentioned GitHub issue.
To install the latest DALI release for the latest CUDA version (12.x):
pip install nvidia-dali-cuda120 # or pip install --extra-index-url https://pypi.nvidia.com --upgrade nvidia-dali-cuda120
DALI requires NVIDIA driver supporting the appropriate CUDA version. In case of DALI based on CUDA 12, it requires CUDA Toolkit to be installed.
DALI comes preinstalled in the TensorFlow, PyTorch, and PaddlePaddle containers on NVIDIA GPU Cloud.
For other installation paths (TensorFlow plugin, older CUDA version, nightly and weekly builds, etc), and specific requirements please refer to the Installation Guide.
To build DALI from source, please refer to the Compilation Guide.
An introduction to DALI can be found in the Getting Started page.
More advanced examples can be found in the Examples and Tutorials page.
For an interactive version (Jupyter notebook) of the examples, go to the docs/examples directory.
Note: Select the Latest Release Documentation or the Nightly Release Documentation, which stays in sync with the main branch, depending on your version.
We welcome contributions to DALI. To contribute to DALI and make pull requests, follow the guidelines outlined in the Contributing document.
If you are looking for a task good for the start please check one from external contribution welcome label.
We appreciate feedback, questions or bug reports. When you need help with the code, follow the process outlined in the Stack Overflow document. Ensure that the posted examples are:
DALI was originally built with major contributions from Trevor Gale, Przemek Tredak, Simon Layton, Andrei Ivanov and Serge Panev.
| Back | FazBrowse Home | New Git URL |