<spanclass="m-breadcrumb"><ahref="cudaStandardAlgorithms.html">CUDA Standard Algorithms</a> »</span>
Parallel Transforms
</h1>
<navclass="m-block m-default">
<h3>Contents</h3>
<ul>
<li><ahref="#CUDASTDParallelTransformsIncludeTheHeader">Include the Header</a></li>
<li><ahref="#CUDASTDTransformARangeOfItems">Transform a Range of Items</a></li>
<li><ahref="#CUDASTDTransformTwoRangesOfItems">Transform Two Ranges of Items</a></li>
</ul>
</nav>
<p>Taskflow provides template methods for transforming ranges of items to different outputs.</p><sectionid="CUDASTDParallelTransformsIncludeTheHeader"><h2><ahref="#CUDASTDParallelTransformsIncludeTheHeader">Include the Header</a></h2><p>You need to include the header file, <code>taskflow/cuda/algorithm/transform.hpp</code>, for using the parallel-transform algorithm.</p><preclass="m-code"><spanclass="cp">#include</span><spanclass="w"></span><spanclass="cpf"><taskflow/cuda/algorithm/transform.hpp></span></pre></section><sectionid="CUDASTDTransformARangeOfItems"><h2><ahref="#CUDASTDTransformARangeOfItems">Transform a Range of Items</a></h2><p>Parallel-transform algorithm applies the given transform function to a range of items and store the result in another range specified by two iterators, <code>first</code> and <code>last</code>. The task created by <ahref="namespacetf.html#a3ed764530620a419e3400e1f9ab6c956" class="m-doc">tf::<wbr/>cuda_transform(P&& p, I first, I last, O output, C op)</a> represents a parallel execution for the following loop:</p><preclass="m-code"><spanclass="k">while</span><spanclass="w"></span><spanclass="p">(</span><spanclass="n">first</span><spanclass="w"></span><spanclass="o">!=</span><spanclass="w"></span><spanclass="n">last</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span>
<spanclass="p">}</span></pre><p>The following example creates a transform kernel that transforms an input range of <code>N</code> items to an output range by multiplying each item by 10.</p><preclass="m-code"><spanclass="n">tf</span><spanclass="o">::</span><spanclass="n">cudaDefaultExecutionPolicy</span><spanclass="w"></span><spanclass="n">policy</span><spanclass="p">;</span>
<spanclass="c1">// synchronize the execution</span>
<spanclass="n">policy</span><spanclass="p">.</span><spanclass="n">synchronize</span><spanclass="p">();</span></pre><p>Each iteration is independent of each other and is assigned one kernel thread to run the callable. The transform algorithm runs <em>asynchronously</em> through the stream specified in the execution policy. You need to synchronize the stream to obtain correct results.</p></section><sectionid="CUDASTDTransformTwoRangesOfItems"><h2><ahref="#CUDASTDTransformTwoRangesOfItems">Transform Two Ranges of Items</a></h2><p>You can transform two ranges of items to an output range through a binary operator. The task created by <ahref="namespacetf.html#abdcb5b755f7ace2aa452541d5bf93b5f" class="m-doc">tf::<wbr/>cuda_transform(P&& p, I1 first1, I1 last1, I2 first2, O output, C op)</a> represents a parallel execution for the following loop:</p><preclass="m-code"><spanclass="k">while</span><spanclass="w"></span><spanclass="p">(</span><spanclass="n">first1</span><spanclass="w"></span><spanclass="o">!=</span><spanclass="w"></span><spanclass="n">last1</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span>
<spanclass="p">}</span></pre><p>The following example creates a transform kernel that transforms two input ranges of <code>N</code> items to an output range by summing each pair of items in the input ranges.</p><preclass="m-code"><spanclass="n">tf</span><spanclass="o">::</span><spanclass="n">cudaDefaultExecutionPolicy</span><spanclass="w"></span><spanclass="n">policy</span><spanclass="p">;</span>