<spanclass="m-breadcrumb"><ahref="cudaStandardAlgorithms.html">CUDA Standard Algorithms</a> »</span>
Parallel Merge
</h1>
<navclass="m-block m-default">
<h3>Contents</h3>
<ul>
<li><ahref="#CUDASTDMergeIncludeTheHeader">Include the Header</a></li>
<li><ahref="#CUDASTDMergeItems">Merge Two Sorted Ranges of Items</a></li>
<li><ahref="#CUDASTDMergeKeyValueItems">Merge Two Sorted Ranges of Key-Value Items</a></li>
</ul>
</nav>
<p>Taskflow provides standalone template methods for merging two sorted ranges of items into a sorted range of items.</p><sectionid="CUDASTDMergeIncludeTheHeader"><h2><ahref="#CUDASTDMergeIncludeTheHeader">Include the Header</a></h2><p>You need to include the header file, <code>taskflow/cuda/algorithm/merge.hpp</code>, for using the parallel-merge algorithm.</p><preclass="m-code"><spanclass="cp">#include</span><spanclass="w"></span><spanclass="cpf"><taskflow/cuda/algorithm/merge.hpp></span><spanclass="cp"></span></pre></section><sectionid="CUDASTDMergeItems"><h2><ahref="#CUDASTDMergeItems">Merge Two Sorted Ranges of Items</a></h2><p><ahref="namespacetf.html#a37ec481149c2f01669353033d75ed72a" class="m-doc">tf::<wbr/>cuda_merge</a> merges two sorted ranges of items into a sorted range. The following code merges two sorted arrays <code>input_1</code> and <code>input_2</code>, each of 1000 items, into a sorted array <code>output</code> of 2000 items.</p><preclass="m-code"><spanclass="k">const</span><spanclass="w"></span><spanclass="kt">size_t</span><spanclass="w"></span><spanclass="n">N</span><spanclass="w"></span><spanclass="o">=</span><spanclass="w"></span><spanclass="mi">1000</span><spanclass="p">;</span><spanclass="w"></span>
<spanclass="n">cudaFree</span><spanclass="p">(</span><spanclass="n">buffer</span><spanclass="p">);</span><spanclass="w"></span></pre><p>The merge algorithm runs <em>asynchronously</em> through the stream specified in the execution policy. You need to synchronize the stream to obtain correct results. Since the GPU merge algorithm may require extra buffer to store the temporary results, you need to provide a buffer of size at least larger or equal to the value returned from <code><ahref="classtf_1_1cudaExecutionPolicy.html#a1febbe549d9cbe4502a5b66167ab9553" class="m-doc">tf::<wbr/>cudaDefaultExecutionPolicy::<wbr/>merge_bufsz</a></code>. The buffer size depends only on the two input vector sizes.</p><asideclass="m-note m-warning"><h4>Attention</h4><p>You must keep the buffer alive before the merge call completes.</p></aside></section><sectionid="CUDASTDMergeKeyValueItems"><h2><ahref="#CUDASTDMergeKeyValueItems">Merge Two Sorted Ranges of Key-Value Items</a></h2><p><ahref="namespacetf.html#aa84d4c68d2cbe9f6efc4a1eb1a115458" class="m-doc">tf::<wbr/>cuda_merge_by_key</a> performs key-value merge over two sorted ranges in a similar way to <ahref="namespacetf.html#a37ec481149c2f01669353033d75ed72a" class="m-doc">tf::<wbr/>cuda_merge</a>; additionally, it copies elements from the two ranges of values associated with the two input keys, respectively. The following code performs key-value merge over <code>a</code> and <code>b:</code></p><preclass="m-code"><spanclass="k">const</span><spanclass="w"></span><spanclass="kt">size_t</span><spanclass="w"></span><spanclass="n">N</span><spanclass="w"></span><spanclass="o">=</span><spanclass="w"></span><spanclass="mi">2</span><spanclass="p">;</span><spanclass="w"></span>
<spanclass="n">cudaFree</span><spanclass="p">(</span><spanclass="n">c_vals</span><spanclass="p">);</span><spanclass="w"></span></pre><p>Since the GPU merge algorithm may require extra buffer to store the temporary results, you need to provide a buffer of size at least larger or equal to the value returned from <code><ahref="classtf_1_1cudaExecutionPolicy.html#a1febbe549d9cbe4502a5b66167ab9553" class="m-doc">tf::<wbr/>cudaDefaultExecutionPolicy::<wbr/>merge_bufsz</a></code>. The buffer size depends only on the two input vector sizes.</p></section>
</div>
</div>
</div>
</article></main>
<divclass="m-doc-search" id="search">
<ahref="#!" onclick="return hideSearch()"></a>
<divclass="m-container">
<divclass="m-row">
<divclass="m-col-m-8 m-push-m-2">
<divclass="m-doc-search-header m-text m-small">
<div><spanclass="m-label m-default">Tab</span> / <spanclass="m-label m-default">T</span> to search, <spanclass="m-label m-default">Esc</span> to close</div>