<p>Taskflow provides template methods for transforming ranges of items to different outputs.</p><sectionid="CUDASTDTransformARangeOfItems"><h2><ahref="#CUDASTDTransformARangeOfItems">Transform a Range of Items</a></h2><p>Parallel-transform algorithm applies the given transform function to a range of items and store the result in another range specified by two iterators, <code>first</code> and <code>last</code>. The task created by <ahref="namespacetf.html#a3ed764530620a419e3400e1f9ab6c956" class="m-doc">tf::<wbr/>cuda_transform(P&& p, I first, I last, O output, C op)</a> represents a parallel execution for the following loop:</p><preclass="m-code"><spanclass="k">while</span><spanclass="p">(</span><spanclass="n">first</span><spanclass="o">!=</span><spanclass="n">last</span><spanclass="p">)</span><spanclass="p">{</span>
<spanclass="p">}</span></pre><p>The following example creates a transform kernel that transforms an input range of <code>N</code> items to an output range by multiplying each item by 10.</p><preclass="m-code"><spanclass="n">tf</span><spanclass="o">::</span><spanclass="n">cuda_transform</span><spanclass="p">(</span>
<spanclass="p">);</span></pre><p>Each iteration is independent of each other and is assigned one kernel thread to run the callable. Since the callable runs on GPU, it must be declared with a <code>__device__</code> specifier.</p></section><sectionid="CUDASTDTransformTwoRangesOfItems"><h2><ahref="#CUDASTDTransformTwoRangesOfItems">Transform Two Ranges of Items</a></h2><p>You can transform two ranges of items to an output range through a binary operator. The task created by <ahref="namespacetf.html#abdcb5b755f7ace2aa452541d5bf93b5f" class="m-doc">tf::<wbr/>cuda_transform(P&& p, I1 first1, I1 last1, I2 first2, O output, C op)</a> represents a parallel execution for the following loop:</p><preclass="m-code"><spanclass="k">while</span><spanclass="p">(</span><spanclass="n">first1</span><spanclass="o">!=</span><spanclass="n">last1</span><spanclass="p">)</span><spanclass="p">{</span>
<spanclass="p">}</span></pre><p>The following example creates a transform kernel that transforms two input ranges of <code>N</code> items to an output range by summing each pair of items in the input ranges.</p><preclass="m-code"><spanclass="c1">// output[i] = input1[i] + inpu2[i]</span>
<spanclass="p">);</span></pre></section><sectionid="CUDASTDInvokeTransformAsynchronously"><h2><ahref="#CUDASTDInvokeTransformAsynchronously">Invoke Parallel-Transform Kernels Asynchronously</a></h2><p>By default, <ahref="namespacetf.html#a3ed764530620a419e3400e1f9ab6c956" class="m-doc">tf::<wbr/>cuda_transform</a> blocks until all kernels finish. You can use <ahref="namespacetf.html#a0505c26296a6d028bc177b8fb8b4451e" class="m-doc">tf::<wbr/>cuda_transform_async</a> to invoke parallel-transform kernels asynchronously and explicitly synchronize them at another place of your program.</p><preclass="m-code"><spanclass="k">auto</span><spanclass="n">p</span><spanclass="o">=</span><spanclass="n">tf</span><spanclass="o">::</span><spanclass="n">cudaDefaultExecutionPolicy</span><spanclass="p">(</span><spanclass="n">my_stream</span><spanclass="p">);</span>