Fixed an out-of-bounds memory access in the VectorAdd example when launching multiple work-groups.
Previously each work-group calculated its starting offset as group_id * N, even though the input and output buffer contained only N elements.
Result - Work-groups with group_id > 0 tried access memory beyond the allocated buffer, which caused UR_RESULT_ERROR_DEVICE_LOST.
Fix - The kernel now uses the global work-item ID and global range to distribute the N elements work-items using a grid stride loop. This ensures the buffers accesses are within the boundary.
Type of change
Please delete options that are not relevant. Add a 'X' to the one that is applicable.
Bug fix (non-breaking change which fixes an issue)
New feature (non-breaking change which adds functionality)
Implement fixes for ONSAM Jiras
How Has This Been Tested?
Please describe the tests that you ran to verify your changes. Provide instructions so we can reproduce. Please also list any relevant details for your test configuration
Command Line
oneapi-cli
Visual Studio
Eclipse IDE
VSCode
When compiling the compliler flag "-Wall -Wformat-security -Werror=format-security" was used
Security and Legal
OSPDT Approval (see Project Manager for assistance)
Compile using the following compiler flags and fix any warnings, the falgs are: "/Wall -Wformat-security -Werror=format-security"
Bandit Scans (Python only)
Virus scan
Review
Review DPC++ code with Paul Peterseon. (GitHub User: pmpeter1)
Review readme with Tom Lenth(@tomlenth) and/or Project Manager
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixed an out-of-bounds memory access in the VectorAdd example when launching multiple work-groups.
Previously each work-group calculated its starting offset as group_id * N, even though the input and output buffer contained only N elements.
Result - Work-groups with group_id > 0 tried access memory beyond the allocated buffer, which caused UR_RESULT_ERROR_DEVICE_LOST.
Fix - The kernel now uses the global work-item ID and global range to distribute the N elements work-items using a grid stride loop. This ensures the buffers accesses are within the boundary.
Type of change
Please delete options that are not relevant. Add a 'X' to the one that is applicable.
How Has This Been Tested?
Please describe the tests that you ran to verify your changes. Provide instructions so we can reproduce. Please also list any relevant details for your test configuration
Security and Legal
Review