<p><ahref="classtf_1_1cudaFlow.html" class="m-doc">tf::<wbr/>cudaFlow</a> provides two template methods, <ahref="classtf_1_1cudaFlow.html#a1a681f6223853b6445dcfdad07e4d0fd" class="m-doc">tf::<wbr/>cudaFlow::<wbr/>for_each</a> and <ahref="classtf_1_1cudaFlow.html#a34f1ea89e5651faa6e8af522a42556ac" class="m-doc">tf::<wbr/>cudaFlow::<wbr/>for_each_index</a>, for creating tasks to perform parallel iterations over a range of items.</p><sectionid="CUDAForEachIncludeTheHeader"><h2><ahref="#CUDAForEachIncludeTheHeader">Include the Header</a></h2><p>You need to include the header file, <code>taskflow/cuda/algorithm/for_each.hpp</code>, for creating a parallel-iteration task.</p><preclass="m-code"><spanclass="cp">#include</span><spanclass="w"></span><spanclass="cpf"><taskflow/cuda/algorithm/for_each.hpp></span></pre></section><sectionid="ForEachCUDAIndexBasedParallelFor"><h2><ahref="#ForEachCUDAIndexBasedParallelFor">Index-based Parallel Iterations</a></h2><p>Index-based parallel-for performs parallel iterations over a range <code>[first, last)</code> with the given <code>step</code> size. The task created by <ahref="classtf_1_1cudaFlow.html#a34f1ea89e5651faa6e8af522a42556ac" class="m-doc">tf::<wbr/>cudaFlow::<wbr/>for_each_index(I first, I last, I step, C callable)</a> represents a kernel of parallel execution for the following loop:</p><preclass="m-code"><spanclass="c1">// positive step: first, first+step, first+2*step, ...</span>
<spanclass="p">}</span></pre><p>Each iteration <code>i</code> is independent of each other and is assigned one kernel thread to run the callable. Since the callable runs on GPU, it must be declared with a <code>__device__</code> specifier. The following example creates a kernel that assigns each entry of <code>gpu_data</code> to 1 over the range [0, 100) with step size 1.</p><preclass="m-code"><spanclass="c1">// assigns each element in gpu_data to 1 over the range [0, 100) with step size 1</span>
<spanclass="p">});</span></pre></section><sectionid="ForEachCUDAIteratorBasedParallelIterations"><h2><ahref="#ForEachCUDAIteratorBasedParallelIterations">Iterator-based Parallel Iterations</a></h2><p>Iterator-based parallel-for performs parallel iterations over a range specified by two STL-styled iterators, <code>first</code> and <code>last</code>. The task created by <ahref="classtf_1_1cudaFlow.html#a1a681f6223853b6445dcfdad07e4d0fd" class="m-doc">tf::<wbr/>cudaFlow::<wbr/>for_each(I first, I last, C callable)</a> represents a parallel execution of the following loop:</p><preclass="m-code"><spanclass="k">for</span><spanclass="p">(</span><spanclass="k">auto</span><spanclass="w"></span><spanclass="n">i</span><spanclass="o">=</span><spanclass="n">first</span><spanclass="p">;</span><spanclass="w"></span><spanclass="n">i</span><spanclass="o"><</span><spanclass="n">last</span><spanclass="p">;</span><spanclass="w"></span><spanclass="n">i</span><spanclass="o">++</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span>
<spanclass="p">}</span></pre><p>The two iterators, <code>first</code> and <code>last</code>, are typically two raw pointers to the first element and the next to the last element in the range in GPU memory space. The following example creates a <code>for_each</code> kernel that assigns each element in <code>gpu_data</code> to 1 over the range <code>[gpu_data, gpu_data + 1000)</code>.</p><preclass="m-code"><spanclass="c1">// assigns each element to 1 over the range [gpu_data, gpu_data + 1000)</span>
<spanclass="p">});</span><spanclass="w"></span></pre><p>Each iteration is independent of each other and is assigned one kernel thread to run the callable. Since the callable runs on GPU, it must be declared with a <code>__device__</code> specifier.</p></section><sectionid="ForEachCUDAMiscellaneousItems"><h2><ahref="#ForEachCUDAMiscellaneousItems">Miscellaneous Items</a></h2><p>The parallel-iteration algorithms are also available in <ahref="classtf_1_1cudaFlowCapturer.html#a0b2f1bcd59f0b42e0f823818348b4ae7" class="m-doc">tf::<wbr/>cudaFlowCapturer::<wbr/>for_each</a> and <ahref="classtf_1_1cudaFlowCapturer.html#aeb877f42ee3a627c40f1c9c84e31ba3c" class="m-doc">tf::<wbr/>cudaFlowCapturer::<wbr/>for_each_index</a>.</p></section>
</div>
</div>
</div>
</article></main>
<divclass="m-doc-search" id="search">
<ahref="#!" onclick="return hideSearch()"></a>
<divclass="m-container">
<divclass="m-row">
<divclass="m-col-m-8 m-push-m-2">
<divclass="m-doc-search-header m-text m-small">
<div><spanclass="m-label m-default">Tab</span> / <spanclass="m-label m-default">T</span> to search, <spanclass="m-label m-default">Esc</span> to close</div>