FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Add ONNX Runtime tensor conversions by yuanknv · Pull Request #19 · ros2/rosidl_buffer_backends · GitHub

Repository navigation

Add ONNX Runtime tensor conversions - #19

Open
yuanknv wants to merge 58 commits into
feature/torch-conversions-pyfrom
feature/onnxruntime-conversions
Open

yuanknv wants to merge 58 commits into
feature/torch-conversions-pyfrom
feature/onnxruntime-conversions

Conversation

yuanknv commented Aug 28, 2026 •
edited
Loading

Copy link
Copy Markdown
Contributor

Add host and accelerator tensor views with inference and inter-process
integration coverage, using runtime-discovered conversion plugins so the same
application code can be deployed with different backend Debian packages.

Dependencies

Depends on ros/rosdistro#53982, which
adds the cuda-toolkit rosdep key for Ubuntu Resolute. That PR must be merged
before this PR.

Description

  1. Framework-native cores and plugins

    • onnxruntime_conversions and onnxruntime_conversions_py provide stable
      C++ and Python APIs backed by separately packaged device plugins. This PR
      includes host and CUDA implementations.
    • The cores discover the installed plugins when a process starts, allowing
      the same application to use the backend packages selected for a deployment.
    • Host memory has priority 0 and the current accelerator plugin has priority
      100. Callers can select a backend per operation.
    • Additional accelerators can be added as new plugin packages without
      changing the framework-facing API or existing application code.
  2. Zero-copy tensor views

    • Each plugin acquires storage from its ROS buffer backend and keeps that
      lease alive for the lifetime of the tensor view.
    • C++ wraps the leased pointer with Ort::Value::CreateTensor, while Python
      exposes the same storage through ONNX Runtime's DLPack API.
    • Copy helpers can also write ONNX Runtime tensors into existing or newly
      allocated messages when a zero-copy view is not the requested operation.
  3. ONNX Runtime providers

    • Debian provider packages bundle ONNX Runtime 1.29.0. Source builds can
      reuse compatible existing installations or use that pinned fallback.
    • CPU-only providers remain available for independent consumers. The CUDA
      plugin requires CUDA Toolkit >=13.1 and a CUDA-enabled provider.
    • Future accelerator plugins can introduce matching provider packages while
      retaining the same conversion APIs and packaging boundary.

Is this user-facing behavior change?

Did you use Generative AI?

Yes. OpenAI Codex (GPT-5) was used to assist with the C++ and Python plugin
refactor, provider packaging, tests, Docker/Debian validation tooling, and
documentation. The resulting code was manually audited and validated using the
repository's C++ and Python tests.

Additional Information

yuanknv self-assigned this Aug 29, 2026
yuanknv force-pushed the feature/onnxruntime-conversions branch from 363fec1 to 8cea712 Compare August 31, 2026 19:54
yuanknv changed the base branch from main to nvcyc/cuda_buffer_py September 1, 2026 20:16
yuanknv force-pushed the feature/onnxruntime-conversions branch from c8e4c37 to 5cf8ff1 Compare September 4, 2026 00:50
yuanknv changed the base branch from nvcyc/cuda_buffer_py to feature/torch-conversions-py September 4, 2026 00:52
yuanknv force-pushed the feature/onnxruntime-conversions branch 3 times, most recently from d5edc0a to 3844f23 Compare September 9, 2026 06:58
Add CPU and optional CUDA tensor views with inference and inter-process integration coverage.
Keep CPU installations platform-neutral while making CUDA runtime and buffer dependencies explicit and deterministic.
Select the runtime provider at source-build time and expose one C++ and Python conversion API across CPU and optional CUDA backends.
Coordinate the Python runtime with the native provider variant and group both conversion packages under one source tree.
Remove the hidden Python stream fallback so applications control stream ownership and consistently share it with ONNX Runtime.
Defer optional CUDA-buffer integration to consumer configuration so the conversion package remains a header-only portable interface.
Rename the conversion adapters to plugins, move the shared API into the
onnxruntime_conversions facade package, and add mutually exclusive CPU and
CUDA runtime packages for C++ and Python. The CUDA plugin now serves both
cpu and cuda backends, so the CUDA runtime no longer ships a second plugin
class.
…k core

onnxruntime_conversions becomes a header-only adapter over
dlpack_conversions, and onnxruntime_conversions_py becomes pure Python
over dlpack_conversions_py. Storage no longer arrives through
ONNX-specific plugins, so the per-device packages are gone:
onnxruntime_conversions_cpu, onnxruntime_conversions_cuda,
onnxruntime_conversions_py_core, onnxruntime_conversions_py_cpu and
onnxruntime_conversions_py_cuda are removed, and their tests move into
the two remaining packages.

Execution provider selection moves into the adapter, where
configure_session_options and session_providers name a provider for the
backend that allocated the storage. Because the adapter compiles in the
consumer's translation unit, an ONNX Runtime upgrade no longer requires
rebuilding the storage plugins, and CPU, CUDA and ROCm storage can be
installed and chosen at runtime.

BREAKING CHANGE: onnxruntime_conversions is now header-only and the
per-device conversion plugin packages no longer exist. Replace
allocate_tensor_msg device arguments with a backend name, and build
session options through configure_session_options or session_providers.
…uildable

The Python to_tensor_msg staged every copy through OrtValue.numpy() and
update_inplace, neither of which works on external device memory, so any
device destination failed outright. Route it through the DLPack core's new
copy path instead, which hands the copy to the storage plugin that owns
the memory.

The pub/sub component nodes call the CUDA runtime directly but were built
unconditionally and never linked cudart, so a CPU-only build of this
package could not compile its tests. Build them, the CUDA gtest, and the
launch test only where the toolkit is present.

Also drop the vendor conflicts against packages this branch deletes, drop
the launch test dependencies the pytest rewrite left behind, and reject a
nonzero DLPack byte_offset rather than silently handing ONNX Runtime a
base pointer.
The READMEs still named the per-device plugin packages this refactor
replaced and still passed an Ort::MemoryInfo and a ConversionConfiguration
that the adapter no longer takes, so every documented example was
uncompilable. Describe the storage plugin split and the current
signatures.
yuanknv marked this pull request as ready for review September 11, 2026 23:58
Accept stable ONNX Runtime >=1.27.0 with provider and API checks. Pin fallback providers to 1.29.0, select CUDA dependencies by toolkit family, and document source reuse and Debian installation.
Share Identity and MatMul fixtures across native unit and interprocess tests. Construct the graphs with ONNX Runtime instead of embedding serialized byte arrays.
Allow toolkit versions >=13.1,<14 and document metapackage selection and linked runtime dependencies.
Stage ARM64 CUDA libraries from the ONNX Runtime wheel and reuse installed
SDKs using version metadata and CUDA provider checks. Remove the native
compile/run probe, Python tensor-operation checks and CUDA runtime hook.

Declare NumPy for Python vendor builds and document source reuse, fallback
versions and existing-toolkit builds in the conversion README.

Validated vendor selection and tests, CPU Debian consumers, CUDA wheel
staging and GPU import orders on x86. ARM source builds were validated
before the final discovery cleanup.
yuanknv force-pushed the feature/onnxruntime-conversions branch from 1a50102 to d30cb5a Compare September 17, 2026 01:48
yuanknv force-pushed the feature/torch-conversions-py branch 2 times, most recently from 4c036ff to cff1b82 Compare October 1, 2026 18:16
yuanknv force-pushed the feature/onnxruntime-conversions branch from 257bfa6 to 8f07685 Compare October 1, 2026 18:16
Restructure Python and C++ examples and document accelerator setup and Debian/source requirements. Remove the ROSIDL_TENSOR_BACKEND override and CUDA upper bound, and configure cppcheck with GoogleTest definitions.
yuanknv force-pushed the feature/torch-conversions-py branch from cff1b82 to 31e3687 Compare October 2, 2026 22:24
yuanknv force-pushed the feature/onnxruntime-conversions branch from 8f07685 to 67b26f8 Compare October 2, 2026 22:24
yuanknv force-pushed the feature/torch-conversions-py branch from 31e3687 to 5023040 Compare October 5, 2026 19:33
yuanknv force-pushed the feature/onnxruntime-conversions branch from 67b26f8 to d582e99 Compare October 5, 2026 19:34
yuanknv force-pushed the feature/torch-conversions-py branch from 5023040 to 31e3687 Compare October 5, 2026 19:57
yuanknv force-pushed the feature/onnxruntime-conversions branch from d582e99 to 67b26f8 Compare October 5, 2026 19:57
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant


Back | FazBrowse Home | New Git URL