FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

[lapack][cuSOLVER] Support multiple RHS in potrs_batch by zjin-lcf · Pull Request #777 · uxlfoundation/oneMath · GitHub

[lapack][cuSOLVER] Support multiple RHS in potrs_batch - #777

Open
zjin-lcf wants to merge 3 commits into
uxlfoundation:developfrom
zjin-lcf:feature/cusolver-potrs-multiple-rhs
Open

zjin-lcf wants to merge 3 commits into
uxlfoundation:developfrom
zjin-lcf:feature/cusolver-potrs-multiple-rhs

Conversation

Copy link
Copy Markdown
Contributor

Summary

  • support nrhs > 1 in buffer-strided, USM-strided, and USM-grouped potrs_batch by issuing one native batched solve per RHS column
  • upload all RHS pointer arrays before native work to avoid asynchronous pointer-array races
  • provide zero-initialized info arrays and release temporary device allocations on success and failure
  • document that native-call cost scales with nrhs

Extracted from #768 for focused review. Depends on #773 for shared batched devInfo reporting; until that PR merges, this branch intentionally includes its foundation commit.

Test plan

  • clang-format 19.1 check
  • cuSOLVER backend and runtime/compile-time LAPACK test binaries compile and link
  • runtime execution on the current host (blocked before test discovery by a pre-existing SYCL CUDA_ERROR_UNKNOWN, also reproduced with the old build)

The same implementation previously passed grouped and strided potrs_batch accuracy tests with multiple right-hand sides on NVIDIA A100 as part of #768.

zjin-lcf and others added 3 commits September 18, 2026 14:51
The USM overload of get_cusolver_devinfo copied a single int regardless of
the number of matrices queried and did not wait on the asynchronous copy, so
getrf_batch silently only checked the first matrix of a batch.

Add lapack_info_check_batch, which collects the info value of every matrix
and throws a lapack::batch_error listing all the failing ones, as done by the
rocSOLVER backend.

Co-authored-by: Cursor <cursoragent@cursor.com>
cusolverDnXpotrsBatched only solves for a single right hand side, but solving
the columns of B one at a time is equivalent, so the strided and group
batches no longer report nrhs > 1 as unimplemented. The pointers of all the
columns are uploaded before the first call because the native calls are not
synchronised and would otherwise race with the device array being rewritten.

The batched solves now also pass a real info array, zero initialised because
cuSOLVER only writes it when a parameter is invalid, and the temporary device
allocations of the strided and group batches are released instead of leaked.

Co-authored-by: Cursor <cursoragent@cursor.com>
Call out that multiple right-hand sides require one native batched solve per column, so the cost scales with nrhs.

Co-authored-by: Cursor <cursoragent@cursor.com>
zjin-lcf requested a review from a team as a code owner September 18, 2026 22:02

This branch has not been deployed

No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant


Back | FazBrowse Home | New Git URL