You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.
Dismiss alert
<divclass="textblock"><p>Taskflow provides a <em>task-parallel</em> pipeline programming framework for you to implement a pipeline algorithm. Pipeline parallelism refers to a parallel execution of multiple data tokens through a linear chain of pipes or stages. Each stage processes the data token sent from the previous stage, applies the given callable to that data token, and then sends the result to the next stage. Multiple data tokens can be processed simultaneously across different stages.</p>
</div><!-- fragment --><h1><aclass="anchor" id="UnderstandPipelineScheduling"></a>
Understand the Pipeline Scheduling Framework</h1>
<p>A <aclass="el" href="classtf_1_1Pipeline.html" title="class to create a pipeline scheduling framework">tf::Pipeline</a> is a composable graph object that allows users to parallelize an application using pipeline parallelism. Unlike conventional pipeline programming frameworks (e.g., Intel TBB), <aclass="el" href="classtf_1_1Pipeline.html" title="class to create a pipeline scheduling framework">tf::Pipeline</a> does not provide any data abstraction but a task-parallel framework for users to customize their data layout when leveraging pipeline parallelism.</p>
<divclass="dotgraph">
<iframescrolling="no" frameborder="0" src="dot_pipeline_our_structure.svg" width="376" height="472"><p><b>This browser is not able to show SVG: try Firefox, Chrome, Safari, or Opera instead.</b></p></iframe></div>
<p>The figure above gives an example of our pipeline scheduling framework. The framework consists of three <em>pipes</em> (serial-parallel-serial stages) and four <em>lines</em> (maximum parallelism), where each line processes at most one data token. A pipeline of three pipes and four lines will propagate each data token through a sequential chain of three pipes and can simultaneously process up to four data tokens at the four lines. Each edge represents a task dependency. For example, the edge from <code>pipe-0</code> to <code>pipe-1</code> in line <code>0</code> represents the task dependency between the first and the second pipes in the first line; the edge from <code>pipe-0</code> in line <code>0</code> to <code>pipe-0</code> in line <code>1</code> represents the task dependency between two adjacent lines when processing two data tokens at the same pipe. Each pipe can be either a <em>serial</em> type or a <em>parallel</em> type, where a serial pipe processes data tokens sequentially and a parallel pipe processes different data tokens simultaneously.</p>
<dlclass="section note"><dt>Note</dt><dd>Due to the nature of pipeline, Taskflow requires the first pipe to be a serial type. The pipeline scheduling algorithm operates in a circular fashion with a factor of line count.</dd></dl>
<p>Taskflow leverages modern C++ and template techniques to strike a balance between the <em>expressiveness</em> and <em>generality</em> in designing the pipeline programming model. In general, there are three steps to create a task-parallel pipeline application:</p>
<oltype="1">
<li>Define the pipeline structure (e.g., pipe type, pipe callable, stopping rule, line count)</li>
<li>Define the data storage and layout, if needed for the application</li>
<li>Define the pipeline taskflow graph using composition</li>
</ol>
<p>The following code creates a pipeline scheduling framework using the example from the previous section. The framework schedules a total of five <em>scheduling tokens</em> labeled from 0 to 4. The first pipe stores the token identifier in a custom data storage, <code>buffer</code>, and each of the rest pipes adds one to the input data from the result of the previous pipe and stores the result into the corresponding line entry in the buffer.</p>
<divclass="ttc" id="aclasstf_1_1Executor_html"><divclass="ttname"><ahref="classtf_1_1Executor.html">tf::Executor</a></div><divclass="ttdoc">class to create an executor</div><divclass="ttdef"><b>Definition</b> executor.hpp:62</div></div>
<divclass="ttc" id="aclasstf_1_1Executor_html_a519777f5783981d534e9e53b99712069"><divclass="ttname"><ahref="classtf_1_1Executor.html#a519777f5783981d534e9e53b99712069">tf::Executor::run</a></div><divclass="ttdeci">tf::Future< void > run(Taskflow &taskflow)</div><divclass="ttdoc">runs a taskflow once</div></div>
<divclass="ttc" id="aclasstf_1_1FlowBuilder_html_ac6f22228d4c2ea2e643c4b0d42c0e92a"><divclass="ttname"><ahref="classtf_1_1FlowBuilder.html#ac6f22228d4c2ea2e643c4b0d42c0e92a">tf::FlowBuilder::composed_of</a></div><divclass="ttdeci">Task composed_of(T &object)</div><divclass="ttdoc">creates a module task for the target object</div><divclass="ttdef"><b>Definition</b> flow_builder.hpp:1831</div></div>
<divclass="ttc" id="aclasstf_1_1Pipe_html"><divclass="ttname"><ahref="classtf_1_1Pipe.html">tf::Pipe</a></div><divclass="ttdoc">class to create a pipe object for a pipeline stage</div><divclass="ttdef"><b>Definition</b> pipeline.hpp:144</div></div>
<divclass="ttc" id="aclasstf_1_1Pipeflow_html"><divclass="ttname"><ahref="classtf_1_1Pipeflow.html">tf::Pipeflow</a></div><divclass="ttdoc">class to create a pipeflow object used by the pipe callable</div><divclass="ttdef"><b>Definition</b> pipeline.hpp:43</div></div>
<divclass="ttc" id="aclasstf_1_1Pipeline_html"><divclass="ttname"><ahref="classtf_1_1Pipeline.html">tf::Pipeline</a></div><divclass="ttdoc">class to create a pipeline scheduling framework</div><divclass="ttdef"><b>Definition</b> pipeline.hpp:307</div></div>
<divclass="ttc" id="aclasstf_1_1Task_html_a08ada0425b490997b6ff7f310107e5e3"><divclass="ttname"><ahref="classtf_1_1Task.html#a08ada0425b490997b6ff7f310107e5e3">tf::Task::name</a></div><divclass="ttdeci">const std::string & name() const</div><divclass="ttdoc">queries the name of the task</div><divclass="ttdef"><b>Definition</b> task.hpp:1435</div></div>
<divclass="ttc" id="aclasstf_1_1Taskflow_html"><divclass="ttname"><ahref="classtf_1_1Taskflow.html">tf::Taskflow</a></div><divclass="ttdoc">class to create a taskflow object</div><divclass="ttdef"><b>Definition</b> taskflow.hpp:64</div></div>
<li>Lines 4-5 define the structure of the pipeline scheduling framework </li>
<li>Line 8 defines the data storage as an one-dimensional array of <code>num_lines</code> integers </li>
<li>Line 12 defines the number of lines in the pipeline </li>
<li>Lines 13-23 define the first serial pipe, which will stop the pipeline scheduling at the fifth token </li>
<li>Lines 25-32 define the second parallel pipe </li>
<li>Lines 34-41 define the third serial pipe </li>
<li>Line 45 defines the pipeline taskflow graph using composition </li>
<li>Line 48 executes the taskflow</li>
</ul>
<p>Taskflow leverages <aclass="el" href="RuntimeTasking.html">Runtime Tasking</a> and <aclass="el" href="ComposableTasking.html">Composable Tasking</a> to implement the pipeline scheduling framework. The taskflow graph of this pipeline example is shown as follows, where 1) one condition task is used to decide which runtime task to run and 2) four runtime tasks are used to schedule tokens at four parallel lines, respectively.</p>
<divclass="dotgraph">
<iframescrolling="no" frameborder="0" src="dot_pipeline_basic_dependency_graph.svg" width="566" height="252"><p><b>This browser is not able to show SVG: try Firefox, Chrome, Safari, or Opera instead.</b></p></iframe></div>
<p>In this example, we customize the data storage, <code>buffer</code>, as an one-dimensional array of 4 integers, since the pipeline structure defines only four parallel lines. Each entry of <code>buffer</code> stores stores the data being processed in the corresponding line. For example, <code>buffer[1]</code> stores the processed data at line <code>1</code>. The following figure shows the data layout of <code>buffer</code>.</p>
<divclass="dotgraph">
<iframescrolling="no" frameborder="0" src="dot_pipeline_memory_layout.svg" width="292" height="143"><p><b>This browser is not able to show SVG: try Firefox, Chrome, Safari, or Opera instead.</b></p></iframe></div>
<dlclass="section note"><dt>Note</dt><dd>In practice, you may need to add padding to the data type of the buffer or align it with the cacheline size to avoid false sharing. If the data type varies at different pipes, you can use <ahref="https://en.cppreference.com/w/cpp/utility/variant">std::variant</a> to store the data types in a uniform storage.</dd></dl>
<p>For each scheduling token, you can use <aclass="el" href="classtf_1_1Pipeflow.html#afee054e6a99965d4b3e36ff903227e6c" title="queries the line identifier of the present token">tf::Pipeflow::line()</a> to get its line identifier and <aclass="el" href="classtf_1_1Pipeflow.html#a4914c1f381a3016e98285b019cf60d6d" title="queries the pipe identifier of the present token">tf::Pipeflow::pipe()</a> to get its pipe identifier. For example, if a scheduling token is at the third pipe of the forth line, <aclass="el" href="classtf_1_1Pipeflow.html#afee054e6a99965d4b3e36ff903227e6c" title="queries the line identifier of the present token">tf::Pipeflow::line()</a> will return <code>3</code> and <aclass="el" href="classtf_1_1Pipeflow.html#a4914c1f381a3016e98285b019cf60d6d" title="queries the pipe identifier of the present token">tf::Pipeflow::pipe()</a> will return <code>2</code> (index starts from 0). To stop the execution of the pipeline, you need to call <aclass="el" href="classtf_1_1Pipeflow.html#a830b7f204cb87fff17e8d424918d9453" title="stops the pipeline scheduling">tf::Pipeflow::stop()</a> at the first pipe. Once the stop signal has been triggered, the pipeline will stop scheduling any new tokens after the callable. As we can see from this example, <aclass="el" href="classtf_1_1Pipeline.html" title="class to create a pipeline scheduling framework">tf::Pipeline</a> gives you the full control to customize your application data on top of a pipeline scheduling framework.</p>
<li>Calling <aclass="el" href="classtf_1_1Pipeflow.html#a830b7f204cb87fff17e8d424918d9453" title="stops the pipeline scheduling">tf::Pipeflow::stop()</a> not at the first pipe has no effect on the pipeline scheduling.</li>
<li>In most cases, std::thread::hardware_concurrency is a good number for line count.</li>
</ol>
</dd></dl>
<p>Our pipeline algorithm schedules tokens in a <em>circular</em> manner, with a factor of <code>num_lines</code>. That is, token <code>t</code> will be processed at line <code>t % num_lines</code>. The following snippet shows one of the possible outputs of this pipeline program:</p>
</div><!-- fragment --><p>There are a total of five tokens running through three pipes. Each pipes prints its input data value, except the first pipe that prints its token identifier. Since the second pipe is a parallel pipe, the output can interleave.</p>
<h1><aclass="anchor" id="ConnectWithTasks"></a>
Connect Pipeline with Other Tasks</h1>
<p>You can connect the pipeline module task with other tasks to create a taskflow application that embeds one or multiple pipeline algorithms. We describe three common examples below:</p><ul>
<li><aclass="el" href="#IterateAPipeline">Example 1: Iterate a Pipeline</a></li>
<li><aclass="el" href="#ConcatenateTwoPipelines">Example 2: Concatenate Two Pipelines</a></li>
<p>This example emulates a data streaming application that iteratively runs a stream of data through a pipeline using conditional tasking. The taskflow graph consists of one pipeline module task and one condition task. The pipeline module task processes a stream of data. The condition task decides the availability of data and reruns the pipeline when the next stream of data becomes available. <br/>
<divclass="line"> 4: <spanclass="keyword">const</span><spanclass="keywordtype">size_t</span> num_lines = 4; <spanclass="comment">// maximum parallelism of the pipeline</span></div>
<divclass="line"> 5: </div>
<divclass="line"> 6: <spanclass="keywordtype">int</span> i = 0, N = 0;</div>
<divclass="line"> 7: <spanclass="comment">// custom data storage</span></div>
<divclass="ttc" id="aclasstf_1_1FlowBuilder_html_a4d52a7fe2814b264846a2085e931652c"><divclass="ttname"><ahref="classtf_1_1FlowBuilder.html#a4d52a7fe2814b264846a2085e931652c">tf::FlowBuilder::emplace</a></div><divclass="ttdeci">Task emplace(C &&callable)</div><divclass="ttdoc">creates a static task</div><divclass="ttdef"><b>Definition</b> flow_builder.hpp:1781</div></div>
<divclass="ttc" id="aclasstf_1_1Task_html_a8c78c453295a553c1c016e4062da8588"><divclass="ttname"><ahref="classtf_1_1Task.html#a8c78c453295a553c1c016e4062da8588">tf::Task::precede</a></div><divclass="ttdeci">Task & precede(Ts &&... tasks)</div><divclass="ttdoc">adds precedence links from this to other tasks</div><divclass="ttdef"><b>Definition</b> task.hpp:1305</div></div>
</div><!-- fragment --><p>Debrief:</p>
<ul>
<li>Lines 4-5 define the structure of the pipeline scheduling framework </li>
<li>Line 8 defines the data storage as an one-dimensional array (<code>num_lines</code> integers) </li>
<li>Line 12 defines the number of lines in the pipeline </li>
<li>Lines 13-23 define the first serial pipe, which will stop the pipeline scheduling when <code>i</code> is <code>5</code><br/>
</li>
<li>Lines 25-32 define the second parallel pipe </li>
<li>Lines 34-41 define the third serial pipe </li>
<li>Lines 44-53 define a condition task which returns 0 when <code>N</code> is less than <code>2</code>, otherwise returns <code>1</code></li>
<li>Line 45 resets variable <code>i</code></li>
<li>Lines 56-57 define the pipeline graph using composition </li>
<li>Lines 58-61 define two static tasks </li>
<li>Line 64-66 define the task dependency </li>
<li>Line 69 executes the taskflow</li>
</ul>
<p>The taskflow graph of this pipeline example is illustrated as follows:</p>
<divclass="dotgraph">
<iframescrolling="no" frameborder="0" src="dot_pipeline_example1.svg" width="643" height="488"><p><b>This browser is not able to show SVG: try Firefox, Chrome, Safari, or Opera instead.</b></p></iframe></div>
<p>The following snippet shows one of the possible outputs:</p>
</div><!-- fragment --><p>The pipeline runs twice as controlled by the condition task <code>conditional</code>. The starting token in the second run of the pipeline is <code>5</code> rather than <code>0</code> because the pipeline keeps a stateful number of tokens. The last token is <code>9</code>, which means the pipeline processes in total <code>10</code> scheduling tokens. The first five tokens (token <code>0</code> to <code>4</code>) are processed in the first run, and the remaining five tokens (token <code>5</code> to <code>9</code>) are processed in the second run. In the condition task, we use <code>N</code> as a decision-making counter to process the next stream of data.</p>
<p>This example demonstrates two concatenated pipelines where a sequence of data tokens run synchronously from one pipeline to another pipeline. The first pipeline task precedes the second pipeline task.</p>
<li>Line 8 defines the data storage (<code>num_lines</code> integers) for pipeline <code>pl_1</code></li>
<li>Line 9 defines the data storage (<code>num_lines</code> integers) for pipeline <code>pl_2</code></li>
<li>Lines 14-24 define the first serial pipe in <code>pl_1</code></li>
<li>Lines 26-33 define the second parallel pipe in <code>pl_1</code></li>
<li>Lines 35-42 define the third serial pipe in <code>pl_1</code></li>
<li>Lines 48-59 define the first serial pipe in <code>pl_2</code> that takes the results of <code>pl_1</code> as inputs </li>
<li>Lines 61-68 define the second parallel pipe in <code>pl_2</code></li>
<li>Lines 70-77 define the third serial pipe in <code>pl_2</code></li>
<li>Lines 81-82 define the pipeline graphs using composition </li>
<li>Line 85 defines the task dependency </li>
<li>Line 88 runs the taskflow</li>
</ul>
<p>The taskflow graph of this pipeline example is illustrated as follows:</p>
<divclass="dotgraph">
<iframescrolling="no" frameborder="0" src="dot_pipeline_example2.svg" width="976" height="252"><p><b>This browser is not able to show SVG: try Firefox, Chrome, Safari, or Opera instead.</b></p></iframe></div>
<p>The following snippet shows one of the possible outputs:</p>
</div><!-- fragment --><p>The output of pipelines <code>pl_1</code> and <code>pl_2</code> can be different from run to run because their second pipes are both parallel types. Due to the task dependency between <code>pipeline_1</code> and <code>pipeline_2</code>, the output of <code>pl_1</code> precedes the output of <code>pl_2</code>.</p>
<li>Line 8 defines the data storage (<code>num_lines</code> integers) for pipeline <code>pl_1</code></li>
<li>Line 9 defines the data storage (<code>num_lines</code> integers) for pipeline <code>pl_2</code></li>
<li>Lines 14-24 define the first serial pipe in <code>pl_1</code></li>
<li>Lines 26-33 define the second parallel pipe in <code>pl_1</code></li>
<li>Lines 35-42 define the third serial pipe in <code>pl_1</code></li>
<li>Lines 48-58 define the first serial pipe in <code>pl_2</code></li>
<li>Lines 60-67 define the second parallel pipe in <code>pl_2</code></li>
<li>Lines 69-76 define the third serial pipe in <code>pl_2</code></li>
<li>Lines 79-82 define the pipeline graphs using composition </li>
<li>Lines 83-84 define a static task. </li>
<li>Line 86 defines the task dependency </li>
<li>Line 88 runs the taskflow</li>
</ul>
<p>The taskflow graph of this pipeline example is illustrated as follows:</p>
<divclass="dotgraph">
<iframescrolling="no" frameborder="0" src="dot_pipeline_example3.svg" width="1139" height="252"><p><b>This browser is not able to show SVG: try Firefox, Chrome, Safari, or Opera instead.</b></p></iframe></div>
<p>The following snippet shows one of the possible outputs:</p>
</div><!-- fragment --><p>Because pipeline <code>pl_1</code> and pipeline <code>pl_2</code> are running in parallel, their outputs may interleave.</p>
<h1><aclass="anchor" id="ResetPipeline"></a>
Reset a Pipeline</h1>
<p>Our pipeline scheduling framework keeps a <em>stateful</em> number of scheduled tokens at each submitted run. You can reset the pipeline to the initial state using <aclass="el" href="classtf_1_1Pipeline.html#a311d874b98de6f0def8a7d869e8d15bd" title="resets the pipeline">tf::Pipeline::reset()</a>, where the number of scheduled tokens will start from zero in the next run. Borrowed from <aclass="el" href="#IterateAPipeline">Example 1: Iterate a Pipeline</a>, the program below resets the pipeline at the second iteration (inside the condition task) so the scheduling token will start from zero in the next run.</p>
</div><!-- fragment --><p>The output can be different from run to run, since the second pipe is a parallel type. At the second iteration from the condition task, we reset the pipeline so the token identifier starts from <code>0</code> rather than <code>5</code>. </p>
</div></div><!-- contents -->
</div><!-- PageDoc -->
</div><!-- doc-content -->
<!-- HTML footer for doxygen 1.13.1-->
<!-- start footer part -->
<divid="nav-path" class="navpath"><!-- id is needed for treeview function! -->