<spanclass="m-breadcrumb"><ahref="cudaStandardAlgorithms.html">CUDA Standard Algorithms</a> »</span>
Parallel Find
</h1>
<navclass="m-block m-default">
<h3>Contents</h3>
<ul>
<li><ahref="#CUDASTDFindIncludeTheHeader">Include the Header</a></li>
<li><ahref="#CUDASTDFindItems">Find an Element in a Range</a></li>
<li><ahref="#CUDASTDFindMinItems">Find the Minimum Element in a Range</a></li>
<li><ahref="#CUDASTDFindMaxItems">Find the Maximum Element in a Range</a></li>
</ul>
</nav>
<p>Taskflow provides standalone template methods for finding elements in the given ranges using GPU.</p><sectionid="CUDASTDFindIncludeTheHeader"><h2><ahref="#CUDASTDFindIncludeTheHeader">Include the Header</a></h2><p>You need to include the header file, <code>taskflow/cuda/algorithm/find.hpp</code>, for using the parallel-find algorithm.</p><preclass="m-code"><spanclass="cp">#include</span><spanclass="w"></span><spanclass="cpf"><taskflow/cuda/algorithm/find.hpp></span></pre></section><sectionid="CUDASTDFindItems"><h2><ahref="#CUDASTDFindItems">Find an Element in a Range</a></h2><p><ahref="namespacetf.html#a5f9dabd7c5d0fa5166cf76d9fa5a038e" class="m-doc">tf::<wbr/>cuda_find_if</a> finds the index of the first element in the range <code>[first, last)</code> that satisfies the given criteria. This is equivalent to the parallel execution of the following loop:</p><preclass="m-code"><spanclass="kt">unsigned</span><spanclass="w"></span><spanclass="n">idx</span><spanclass="w"></span><spanclass="o">=</span><spanclass="w"></span><spanclass="mi">0</span><spanclass="p">;</span>
<spanclass="k">return</span><spanclass="w"></span><spanclass="n">idx</span><spanclass="p">;</span></pre><p>If no such an element is found, the size of the range is returned. The following code finds the index of the first element that is dividable by <code>17</code> over a range of one million elements.</p><preclass="m-code"><spanclass="k">const</span><spanclass="w"></span><spanclass="kt">size_t</span><spanclass="w"></span><spanclass="n">N</span><spanclass="w"></span><spanclass="o">=</span><spanclass="w"></span><spanclass="mi">1000000</span><spanclass="p">;</span>
<spanclass="n">cudaFree</span><spanclass="p">(</span><spanclass="n">idx</span><spanclass="p">);</span></pre><p>The find-if algorithm runs <em>asynchronously</em> through the stream specified in the execution policy. You need to synchronize the stream to obtain the correct result.</p></section><sectionid="CUDASTDFindMinItems"><h2><ahref="#CUDASTDFindMinItems">Find the Minimum Element in a Range</a></h2><p><ahref="namespacetf.html#a572c13198191c46765264f8afabe2e9f" class="m-doc">tf::<wbr/>cuda_min_element</a> finds the index of the minimum element in the given range <code>[first, last)</code> using the given comparison function object. This is equivalent to a parallel execution of the following loop:</p><preclass="m-code"><spanclass="k">if</span><spanclass="p">(</span><spanclass="n">first</span><spanclass="w"></span><spanclass="o">==</span><spanclass="w"></span><spanclass="n">last</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span>
<spanclass="k">return</span><spanclass="w"></span><spanclass="n">std</span><spanclass="o">::</span><spanclass="n">distance</span><spanclass="p">(</span><spanclass="n">first</span><spanclass="p">,</span><spanclass="w"></span><spanclass="n">smallest</span><spanclass="p">);</span></pre><p>The following code finds the index of the minimum element in a range of one millions elements using GPU computing:</p><preclass="m-code"><spanclass="k">const</span><spanclass="w"></span><spanclass="kt">size_t</span><spanclass="w"></span><spanclass="n">N</span><spanclass="w"></span><spanclass="o">=</span><spanclass="w"></span><spanclass="mi">1000000</span><spanclass="p">;</span>
<spanclass="n">cudaFree</span><spanclass="p">(</span><spanclass="n">buffer</span><spanclass="p">);</span></pre><p>Since the GPU min-element algorithm may require extra buffer to store the temporary results, you need to provide a buffer of size at least larger or equal to the value returned from <code><ahref="classtf_1_1cudaExecutionPolicy.html#abcafb001cd68c1135392f4bcda5a2a05" class="m-doc">tf::<wbr/>cudaDefaultExecutionPolicy::<wbr/>min_element_bufsz</a></code>.</p><asideclass="m-note m-warning"><h4>Attention</h4><p>You must keep the buffer alive before the <ahref="namespacetf.html#a572c13198191c46765264f8afabe2e9f" class="m-doc">tf::<wbr/>cuda_min_element</a> completes.</p></aside></section><sectionid="CUDASTDFindMaxItems"><h2><ahref="#CUDASTDFindMaxItems">Find the Maximum Element in a Range</a></h2><p>Similar to <ahref="namespacetf.html#a572c13198191c46765264f8afabe2e9f" class="m-doc">tf::<wbr/>cuda_min_element</a>, <ahref="namespacetf.html#a3fc577fd0a8f127770bcf68bc56c073e" class="m-doc">tf::<wbr/>cuda_max_element</a> finds the index of the maximum element in the given range <code>[first, last)</code> using the given comparison function object. This is equivalent to a parallel execution of the following loop:</p><preclass="m-code"><spanclass="k">if</span><spanclass="p">(</span><spanclass="n">first</span><spanclass="w"></span><spanclass="o">==</span><spanclass="w"></span><spanclass="n">last</span><spanclass="p">)</span><spanclass="w"></span><spanclass="p">{</span>
<spanclass="k">return</span><spanclass="w"></span><spanclass="n">std</span><spanclass="o">::</span><spanclass="n">distance</span><spanclass="p">(</span><spanclass="n">first</span><spanclass="p">,</span><spanclass="w"></span><spanclass="n">largest</span><spanclass="p">);</span></pre><p>The following code finds the index of the maximum element in a range of one millions elements using GPU computing:</p><preclass="m-code"><spanclass="k">const</span><spanclass="w"></span><spanclass="kt">size_t</span><spanclass="w"></span><spanclass="n">N</span><spanclass="w"></span><spanclass="o">=</span><spanclass="w"></span><spanclass="mi">1000000</span><spanclass="p">;</span>
<spanclass="n">cudaFree</span><spanclass="p">(</span><spanclass="n">buffer</span><spanclass="p">);</span></pre><p>Since the GPU max-element algorithm may require extra buffer to store the temporary results, you need to provide a buffer of size at least larger or equal to the value returned from <code><ahref="classtf_1_1cudaExecutionPolicy.html#a31fe75c4b0765df3035e12be49af88aa" class="m-doc">tf::<wbr/>cudaDefaultExecutionPolicy::<wbr/>max_element_bufsz</a></code>.</p><asideclass="m-note m-warning"><h4>Attention</h4><p>You must keep the buffer alive before <ahref="namespacetf.html#a3fc577fd0a8f127770bcf68bc56c073e" class="m-doc">tf::<wbr/>cuda_max_element</a> completes.</p></aside></section>
</div>
</div>
</div>
</article></main>
<divclass="m-doc-search" id="search">
<ahref="#!" onclick="return hideSearch()"></a>
<divclass="m-container">
<divclass="m-row">
<divclass="m-col-m-8 m-push-m-2">
<divclass="m-doc-search-header m-text m-small">
<div><spanclass="m-label m-default">Tab</span> / <spanclass="m-label m-default">T</span> to search, <spanclass="m-label m-default">Esc</span> to close</div>