<para><refrefid="classtf_1_1cudaFlow"kindref="compound">tf::cudaFlow</ref> provides template methods for transforming ranges of items to different outputs.</para>
<para>You need to include the header file, <computeroutput>taskflow/cuda/algorithm/transform.hpp</computeroutput>, for creating a parallel-transform task.</para>
<para>Iterator-based parallel-transform applies the given transform function to a range of items and store the result in another range specified by two iterators, <computeroutput>first</computeroutput> and <computeroutput>last</computeroutput>. The task created by <refrefid="classtf_1_1cudaFlow_1af89a9bda182272462a0eda2581536cd8"kindref="member">tf::cudaFlow::transform(I first, I last, O output, C op)</ref> represents a parallel execution for the following loop:</para>
<para>The following example creates a transform kernel that transforms an input range of <computeroutput>N</computeroutput> items to an output range by multiplying each item by 10.</para>
<para>Each iteration is independent of each other and is assigned one kernel thread to run the callable. Since the callable runs on GPU, it must be declared with a <computeroutput>__device__</computeroutput> specifier.</para>
<para>You can transform two ranges of items to an output range through a binary operator. The task created by <refrefid="classtf_1_1cudaFlow_1abab2bfdfc86ef3a764ece4743fdede76"kindref="member">tf::cudaFlow::transform(I1 first1, I1 last1, I2 first2, O output, C op)</ref> represents a parallel execution for the following loop:</para>
<para>The following example creates a transform kernel that transforms two input ranges of <computeroutput>N</computeroutput> items to an output range by summing each pair of items in the input ranges.</para>
<para>The parallel-transform algorithms are also available in <refrefid="classtf_1_1cudaFlowCapturer"kindref="compound">tf::cudaFlowCapturer</ref>. </para>