<para>This chapters discusses how to create a task graph dynamically using asynchronous tasks, which is extremely beneficial for workloads that want to (1) explore task graph parallelism out of dynamic control flow or (2) overlap task graph creation time with individual task execution time. We recommend that you first read <refrefid="AsyncTasking"kindref="compound">Asynchronous Tasking</ref> before digesting this chapter.</para>
<para>When the construct-and-run model of a task graph is not possible in your application, you can use <ref refid="classtf_1_1Executor_1aee02b63d3a91ad5ca5a1c0e71f3e128f" kindref="member">tf::Executor::dependent_async</ref> and <ref refid="classtf_1_1Executor_1a0e2d792f28136b8227b413d0c27d5c7f" kindref="member">tf::Executor::silent_dependent_async</ref> to create a task graph dynamically. This type of parallelism is also known as <emphasis>on-the-fly</emphasis> task graph parallelism, which offers great flexibility for expressing dynamic task graph parallelism. The example below dynamically creates a task graph of four dependent-async tasks, <computeroutput>A</computeroutput>, <computeroutput>B</computeroutput>, <computeroutput>C</computeroutput>, and <computeroutput>D</computeroutput>, where <computeroutput>A</computeroutput> runs before <computeroutput>B</computeroutput> and <computeroutput>C</computeroutput> and <computeroutput>D</computeroutput> runs after <computeroutput>B</computeroutput> and <computeroutput>C:</computeroutput> </para>
<para>Both <refrefid="classtf_1_1Executor_1aee02b63d3a91ad5ca5a1c0e71f3e128f"kindref="member">tf::Executor::dependent_async</ref> and <refrefid="classtf_1_1Executor_1a0e2d792f28136b8227b413d0c27d5c7f"kindref="member">tf::Executor::silent_dependent_async</ref> create a task of type <refrefid="classtf_1_1AsyncTask"kindref="compound">tf::AsyncTask</ref> to run the given function asynchronously. Additionally, <refrefid="classtf_1_1Executor_1aee02b63d3a91ad5ca5a1c0e71f3e128f"kindref="member">tf::Executor::dependent_async</ref> returns a <ulinkurl="https://en.cppreference.com/w/cpp/thread/future">std::future</ulink> that eventually holds the result of the execution. When returning from both calls, the executor has scheduled a worker to run the task whenever its dependencies are met. That is, task execution happens <emphasis>simultaneously</emphasis> with the creation of the task graph, which is different from constructing a Taskflow and running it from an executor, illustrated in the figure below:</para>
<para>Since this model only allows relating a dependency from the current task to a previously created task, you need a correct topological order of graph expression. In our example, there are only two possible topological orderings, either <computeroutput>ABCD</computeroutput> or <computeroutput>ACBD</computeroutput>. The code below shows another feasible order of expressing this dynamic task graph parallelism:</para>
<para>In addition to using <ulinkurl="https://en.cppreference.com/w/cpp/thread/future">std::future</ulink> to synchronize the execution, you can use <refrefid="classtf_1_1Executor_1ab9aa252f70e9a40020a1e5a89d485b85"kindref="member">tf::Executor::wait_for_all</ref> to wait for all scheduled tasks to finish:</para>
<title>Specify a Range of Dependent Async Tasks</title>
<para>Both <refrefid="classtf_1_1Executor_1aee02b63d3a91ad5ca5a1c0e71f3e128f"kindref="member">tf::Executor::dependent_async(F&& func, Tasks&&... tasks)</ref> and <refrefid="classtf_1_1Executor_1a0e2d792f28136b8227b413d0c27d5c7f"kindref="member">tf::Executor::silent_dependent_async(F&& func, Tasks&&... tasks)</ref> accept an arbitrary number of tasks in the dependency list. If the number of task dependencies (i.e., predecessors) is unknown at programming time, such as those relying on runtime variables, you can use the following two overloads to specify predecessor tasks in an iterable range <computeroutput>[first, last)</computeroutput>:</para>
<para><itemizedlist>
<listitem><para><refrefid="classtf_1_1Executor_1a01e51e564f5def845506bcf6b4bb1664"kindref="member">tf::Executor::dependent_async(F&& func, I first, I last)</ref></para>
</listitem><listitem><para><refrefid="classtf_1_1Executor_1aa9b08e47e68ae1e568f18aa7104cb9b1"kindref="member">tf::Executor::silent_dependent_async(F&& func, I first, I last)</ref></para>
</listitem></itemizedlist>
</para>
<para>The code below creates an asynchronous task that depends on <computeroutput>N</computeroutput> previously created asynchronous tasks stored in a vector, where <computeroutput>N</computeroutput> is a runtime variable:</para>
<title>Understand the Lifetime of a Dependent Async Task</title>
<para>A <refrefid="classtf_1_1AsyncTask"kindref="compound">tf::AsyncTask</ref> is a lightweight handle that retains <emphasis>shared</emphasis> ownership of a dependent-async task created by an executor. This shared ownership ensures that the async task remains alive when adding it to the dependency list of another async task, thus avoiding the classical <ulinkurl="https://en.wikipedia.org/wiki/ABA_problem">ABA problem</ulink>.</para>
<para>Currently, <refrefid="classtf_1_1AsyncTask"kindref="compound">tf::AsyncTask</ref> is implemented based on the logic of C++ smart pointer <refrefid="cpp/memory/shared_ptr"kindref="compound"external="/home/thuang295/Code/taskflow/doxygen/cppreference-doxygen-web.tag.xml">std::shared_ptr</ref> and is considered cheap to copy or move as long as only a handful of objects own it. When a worker completes an async task, it will remove the task from the executor, decrementing the number of shared owners by one. If that counter reaches zero, the task is destroyed.</para>
<title>Create a Dynamic Task Graph by Multiple Threads</title>
<para>You can use multiple threads to create a dynamic task graph as long as the order of simultaneously creating tasks is topologically correct. The example below uses creates a dynamic task graph using three threads (including the main thread), where task <computeroutput>A</computeroutput> runs before task <computeroutput>B</computeroutput> and task <computeroutput>C:</computeroutput> </para>
<para>Regardless of <computeroutput>t1</computeroutput> runs before or after <computeroutput>t2</computeroutput>, the resulting topological order is always correct with the graph definition, either <computeroutput>ABC</computeroutput> or <computeroutput>ACB</computeroutput>.</para>
<title>Query the Completion Status of Dependent Async Tasks</title>
<para>When you create a dependent-async task, you can query its completion status by <refrefid="classtf_1_1AsyncTask_1aefeefa30d7cafdfbb7dc8def542e8e51"kindref="member">tf::AsyncTask::is_done</ref>, which returns <computeroutput>true</computeroutput> upon completion or <computeroutput>false</computeroutput> otherwise. A completed dependent-async task indicates that a worker has executed its associated callable.</para>
<para><refrefid="classtf_1_1AsyncTask_1aefeefa30d7cafdfbb7dc8def542e8e51"kindref="member">tf::AsyncTask::is_done</ref> is useful when you need to wait on the result of a dependent-async task before moving onto the next program instruction. Often, <refrefid="classtf_1_1AsyncTask"kindref="compound">tf::AsyncTask</ref> is used together with <refrefid="classtf_1_1Executor_1a0fc6eb19f168dc4a9cd0a7c6187c1d2d"kindref="member">tf::Executor::corun_until</ref> to keep a worker awake in its work-stealing loop to avoid deadlock (see <refrefid="ExecuteTaskflow_1ExecuteATaskflowFromAnInternalWorker"kindref="member">Execute a Taskflow from an Internal Worker</ref> for more details). For instance, the code below implements the famous Fibonacci sequence using recursive asynchronous tasking:</para>