| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
The USM overload of get_cusolver_devinfo copied a single int regardless of the number of matrices queried and did not wait on the asynchronous copy, so getrf_batch silently only checked the first matrix of a batch. Add lapack_info_check_batch, which collects the info value of every matrix and throws a lapack::batch_error listing all the failing ones, as done by the rocSOLVER backend. Co-authored-by: Cursor <cursoragent@cursor.com>
cusolverDnXpotrsBatched only solves for a single right hand side, but solving the columns of B one at a time is equivalent, so the strided and group batches no longer report nrhs > 1 as unimplemented. The pointers of all the columns are uploaded before the first call because the native calls are not synchronised and would otherwise race with the device array being rewritten. The batched solves now also pass a real info array, zero initialised because cuSOLVER only writes it when a parameter is invalid, and the temporary device allocations of the strided and group batches are released instead of leaked. Co-authored-by: Cursor <cursoragent@cursor.com>
Call out that multiple right-hand sides require one native batched solve per column, so the cost scales with nrhs. Co-authored-by: Cursor <cursoragent@cursor.com>
| Back | FazBrowse Home | New Git URL |
Summary
Extracted from #768 for focused review. Depends on #773 for shared batched devInfo reporting; until that PR merges, this branch intentionally includes its foundation commit.
Test plan
The same implementation previously passed grouped and strided potrs_batch accuracy tests with multiple right-hand sides on NVIDIA A100 as part of #768.