| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
On Intel GPUs af::moments returned 0 for every moment slot after the first whenever more than one moment was requested, so AF_MOMENT_FIRST_ORDER came back as (M00, M01, 0, 0). The kernel accumulated into local memory through a hand-rolled float compare-and-swap loop, which Intel's GPU compiler mishandles for the second and later slots while the same code is fine on the Intel CPU OpenCL device. Each work-item now accumulates its rows privately and the group reduces with a tree in local memory, which needs no local atomics and is cheaper. The cross-group global atomic add is unchanged.
…EADS to the launch
| Back | FazBrowse Home | New Git URL |
On Intel GPUs af::moments returned 0 for every moment slot after the first whenever more than one moment was requested, so AF_MOMENT_FIRST_ORDER came back as (M00, M01, 0, 0) and six of the seven moments tests failed on an Arc B580 (driver 32.0.101.8991) and a UHD 770 (32.0.101.7088); the Intel CPU OpenCL device was correct. The kernel accumulated into local memory through a hand-rolled float compare-and-swap loop, and on those two devices that loop returns zeros for the second and later slots; I have no reduced repro or upstream report for why. Each work-item now accumulates its rows privately and the group reduces with a tree in local memory, so the kernel needs no local atomics; the cross-group global atomic add is unchanged. The moments suite passes on both Intel GPUs and the CPU device with this. Not run on AMD or NVIDIA OpenCL.
sift_nonfree.cl carries the same local float compare-and-swap and is left alone here; it needs its own issue.