<para>Taskflow provides standard template methods for scanning a range of items on a CUDA GPU.</para><sect1id="CUDASTDScan_1CUDASTDScanItems">
<title>Scan a Range of Items</title>
<para><refrefid="namespacetf_1affd3cc9cc8550ffead1f63616f22b64f"kindref="member">tf::cuda_inclusive_scan</ref> computes an inclusive prefix sum operation using the given binary operator over a range of elements specified by <computeroutput>[first, last)</computeroutput>. The term "inclusive" means that the i-th input element is included in the i-th sum. The following code computes the inclusive prefix sum over an input array and stores the result in an output array.</para><para><programlistingfilename=".cpp"><codeline><highlightclass="keyword">const</highlight><highlightclass="normal"><sp/></highlight><highlightclass="keywordtype">size_t</highlight><highlightclass="normal"><sp/>N<sp/>=<sp/>1000000;</highlight></codeline>
</programlisting></para><para>On the other hand, <refrefid="namespacetf_1a79efb12473e146c1fddc448e8bb42bd4"kindref="member">tf::cuda_exclusive_scan</ref> computes an exclusive prefix sum operation. The term "exclusive" means that the i-th input element is <emphasis>NOT</emphasis> included in the i-th sum.</para><para><programlistingfilename=".cpp"><codeline><highlightclass="comment">//<sp/>computes<sp/>exclusive<sp/>scan<sp/>over<sp/>input<sp/>and<sp/>stores<sp/>the<sp/>result<sp/>in<sp/>output</highlight><highlightclass="normal"></highlight></codeline>
<para><refrefid="namespacetf_1a98feffc08d80abcabb7e2193fef2b380"kindref="member">tf::cuda_transform_inclusive_scan</ref> transforms each item in the range <computeroutput>[first, last)</computeroutput> and computes an inclusive prefix sum over these transformed items. The following code multiplies each item by 10 and then compute the inclusive prefix sum over 1000000 transformed items.</para><para><programlistingfilename=".cpp"><codeline><highlightclass="keyword">const</highlight><highlightclass="normal"><sp/></highlight><highlightclass="keywordtype">size_t</highlight><highlightclass="normal"><sp/>N<sp/>=<sp/>1000000;</highlight></codeline>
</programlisting></para><para>Similarly, <refrefid="namespacetf_1a4c6f39904ef71525825aeabc0665d806"kindref="member">tf::cuda_transform_exclusive_scan</ref> performs an exclusive prefix sum over a range of transformed items. The following code computes the exclusive prefix sum over 1000000 transformed items each multipled by 10.</para><para><programlistingfilename=".cpp"><codeline><highlightclass="keyword">const</highlight><highlightclass="normal"><sp/></highlight><highlightclass="keywordtype">size_t</highlight><highlightclass="normal"><sp/>N<sp/>=<sp/>1000000;</highlight></codeline>
<para>By default, scan functions block until all kernels finish. You can invoke these kernels asynchronously and explicitly synchronize them at another place of your program. Since our scan kernels rely on additional device memory, you need to provide a buffer of size <emphasis>in bytes</emphasis> equal to (or larger than) the value returned by <refrefid="namespacetf_1afd7e7a0886352201cced5c55205f3229"kindref="member">tf::cuda_scan_buffer_size</ref>.</para><para><programlistingfilename=".cpp"><codeline><highlightclass="keyword">const</highlight><highlightclass="normal"><sp/></highlight><highlightclass="keywordtype">size_t</highlight><highlightclass="normal"><sp/>N<sp/>=<sp/>1000000;</highlight></codeline>
</programlisting></para><para>The allocated buffer must remain alive until the asynchronous call completes.</para><para><simplesectkind="note"><para>The stream given to an asynchronous scan function can be enabled to the capture mode to capture internal kernels into a CUDA graph. </para></simplesect>