<p><ahref="classtf_1_1cudaFlow.html" class="m-doc">tf::<wbr/>cudaFlow</a> provides template methods for transforming ranges of items to different outputs.</p><sectionid="CUDAParallelTransformsIncludeTheHeader"><h2><ahref="#CUDAParallelTransformsIncludeTheHeader">Include the Header</a></h2><p>You need to include the header file, <code>taskflow/cuda/algorithm/transform.hpp</code>, for creating a parallel-transform task.</p><preclass="m-code"><spanclass="cp">#include</span><spanclass="w"></span><spanclass="cpf"><taskflow/cuda/algorithm/transform.hpp></span><spanclass="cp"></span></pre></section><sectionid="cudaFlowTransformARangeOfItems"><h2><ahref="#cudaFlowTransformARangeOfItems">Transform a Range of Items</a></h2><p>Iterator-based parallel-transform applies the given transform function to a range of items and store the result in another range specified by two iterators, <code>first</code> and <code>last</code>. The task created by <ahref="classtf_1_1cudaFlow.html#af89a9bda182272462a0eda2581536cd8" class="m-doc">tf::<wbr/>cudaFlow::<wbr/>transform(I first, I last, O output, C op)</a> represents a parallel execution for the following loop:</p><preclass="m-code"><spanclass="k">while</span><spanclass="w"></span><spanclass="p">(</span><spanclass="n">first</span><spanclass="w"></span><spanclass="o">!=</span><spanclass="w"></span><spanclass="n">last</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span><spanclass="w"></span>
<spanclass="p">}</span><spanclass="w"></span></pre><p>The following example creates a transform kernel that transforms an input range of <code>N</code> items to an output range by multiplying each item by 10.</p><preclass="m-code"><spanclass="c1">// output[i] = input[i] * 10</span>
<spanclass="p">);</span><spanclass="w"></span></pre><p>Each iteration is independent of each other and is assigned one kernel thread to run the callable. Since the callable runs on GPU, it must be declared with a <code>__device__</code> specifier.</p></section><sectionid="cudaFlowTransformTwoRangesOfItems"><h2><ahref="#cudaFlowTransformTwoRangesOfItems">Transform Two Ranges of Items</a></h2><p>You can transform two ranges of items to an output range through a binary operator. The task created by <ahref="classtf_1_1cudaFlow.html#abab2bfdfc86ef3a764ece4743fdede76" class="m-doc">tf::<wbr/>cudaFlow::<wbr/>transform(I1 first1, I1 last1, I2 first2, O output, C op)</a> represents a parallel execution for the following loop:</p><preclass="m-code"><spanclass="k">while</span><spanclass="w"></span><spanclass="p">(</span><spanclass="n">first1</span><spanclass="w"></span><spanclass="o">!=</span><spanclass="w"></span><spanclass="n">last1</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span><spanclass="w"></span>
<spanclass="p">}</span><spanclass="w"></span></pre><p>The following example creates a transform kernel that transforms two input ranges of <code>N</code> items to an output range by summing each pair of items in the input ranges.</p><preclass="m-code"><spanclass="c1">// output[i] = input1[i] + inpu2[i]</span>
<spanclass="p">);</span><spanclass="w"></span></pre></section><sectionid="ParallelTransformCUDAMiscellaneousItems"><h2><ahref="#ParallelTransformCUDAMiscellaneousItems">Miscellaneous Items</a></h2><p>The parallel-transform algorithms are also available in <ahref="classtf_1_1cudaFlowCapturer.html" class="m-doc">tf::<wbr/>cudaFlowCapturer</a>.</p></section>
</div>
</div>
</div>
</article></main>
<divclass="m-doc-search" id="search">
<ahref="#!" onclick="return hideSearch()"></a>
<divclass="m-container">
<divclass="m-row">
<divclass="m-col-m-8 m-push-m-2">
<divclass="m-doc-search-header m-text m-small">
<div><spanclass="m-label m-default">Tab</span> / <spanclass="m-label m-default">T</span> to search, <spanclass="m-label m-default">Esc</span> to close</div>