<sectionid="InstallSYCLCompiler"><h2><ahref="#InstallSYCLCompiler">Install SYCL Compiler</a></h2><p>To compile Taskflow with SYCL code, you need the DPC++ clang compiler, which can be acquired from <ahref="https://intel.github.io/llvm-docs/GetStartedGuide.html">Getting Started with oneAPI DPC++</a>.</p></section><sectionid="CompileTaskflowWithSYCLDirectly"><h2><ahref="#CompileTaskflowWithSYCLDirectly">Compile Source Code Directly</a></h2><p>Taskflow's GPU programming interface for SYCL is <ahref="classtf_1_1syclFlow.html" class="m-doc">tf::<wbr/>syclFlow</a>. Consider the following <code>simple.cpp</code> program that performs the canonical saxpy (single-precision AX + Y) operation on a GPU:</p><preclass="m-code"><spanclass="cp">#include</span><spanclass="cpf"><taskflow/taskflow.hpp></span><spanclass="c1"> // core taskflow routines</span><spanclass="cp"></span>
<spanclass="p">}</span></pre><p>Use DPC++ clang to compile the program with the following options:</p><ul><li><code>-fsycl</code>: enable SYCL compilation mode</li><li><code>-fsycl-targets=nvptx64-nvidia-cuda-sycldevice</code>: enable CUDA target</li><li><code>-fsycl-unnamed-lambda</code>: enable unnamed SYCL lambda kernel</li></ul><preclass="m-console"><spanclass="go">~$ clang++ -fsycl -fsycl-unnamed-lambda \</span>
<spanclass="go"> -fsycl-targets=nvptx64-nvidia-cuda-sycldevice \ # for CUDA target</span>
<spanclass="go">~$ ./simple</span></pre><asideclass="m-note m-warning"><h4>Attention</h4><p>You need to include <code><ahref="syclflow_8hpp.html" class="m-doc">taskflow/<wbr/>syclflow.hpp</a></code> in order to use <ahref="classtf_1_1syclFlow.html" class="m-doc">tf::<wbr/>syclFlow</a>.</p></aside></section><sectionid="CompileTaskflowWithSYCLSeparately"><h2><ahref="#CompileTaskflowWithSYCLSeparately">Compile Source Code Separately</a></h2><p>Large GPU applications often compile a program into separate objects and link them together to form an executable or a library. You can compile your SYCL code into separate object files and link them to form the final executable. Consider the following example that defines two tasks on two different pieces (<code>main.cpp</code> and <code>syclflow.cpp</code>) of source code:</p><preclass="m-code"><spanclass="c1">// main.cpp</span>
<spanclass="n">tf</span><spanclass="o">::</span><spanclass="n">Task</span><spanclass="n">make_syclflow</span><spanclass="p">(</span><spanclass="n">tf</span><spanclass="o">::</span><spanclass="n">Taskflow</span><spanclass="o">&</span><spanclass="n">taskflow</span><spanclass="p">);</span><spanclass="c1">// create a syclFlow task</span>
<spanclass="kr">inline</span><spanclass="n">sycl</span><spanclass="o">::</span><spanclass="n">queue</span><spanclass="n">queue</span><spanclass="p">;</span><spanclass="c1">// create a global sycl queue</span>
<spanclass="p">}</span></pre><p>Compile each source to an object using DPC++ clang:</p><preclass="m-console"><spanclass="go">~$ clang++ -I path/to/taskflow/ -pthread -std=c++17 -c main.cpp -o main.o</span>
<spanclass="gp">#</span> now we have the two compiled .o objects, main.o and syclflow.o
<spanclass="go">~$ ls</span>
<spanclass="go">main.o syclflow.o </span></pre><p>Next, link the two object files to the final executable:</p><preclass="m-console"><spanclass="go">~$ clang++ -fsycl -fsycl-unnamed-lambda \</span>
<spanclass="go"> -fsycl-targets=nvptx64-nvidia-cuda-sycldevice \ # for CUDA target</span>