| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Nekbone solves a standard Poisson equation using a conjugate gradient iteration with a simple or spectral element multigrid preconditioner on a block or linear geometry. It exposes the principal computational kernel to reveal the essential elements of the algorithmic- architectural coupling that is pertinent to Nek5000.
The CUDA/OpenACC branch contains GPU implementations of the conjugate gradient solver. This includes a pure OpenACC implementation as well as a hybrid OpenACC/CUDA implementation with a CUDA kernel for matrix-vector multiplication, based on previous works [1][2]. This implementation can also be compiled with without OpenACC or CUDA. These implementations were tested with the nek_gpu1 example in tests/nek_gpu1. They were also tested with the PGI compilers.
The nek_gpu1 example contains three makenek scripts to compile Nekbone with various configurations:
The different build configurations are determined by the presence/absence of the -acc and -Mcuda flags, as demonstrated in the scripts. These scripts will work without out-of-the-box provided that:
If these defaults are not suitable, you may manually edit the relevant settings in script. You may also compile without MPI by setting the IFMPI option in makenek.
To use the compilation scripts (for example, with OpenACC only):
$ cd test/nek_gpu1 $ ./makenek.acc clean $ ./makenek.acc
The clean step is necessary if you are switching between MPI, OpenACC, and OpenACC/CUDA builds.
Successful compilation will produce an executable called nekbone. You may run the executable through an MPI process manager, for example:
$ mpiexec -n 1 ./nekbone data
The data argument specifies that you will be using the runtime settings from data.rea. See USERGUIDE.pdf for a detailed description of the .rea file format.
The tests/nek_gpu1/scripts/ folder contains simple scripts for comparing results from the different implementations:
To use them:
$ cd test/nek_gpu1 $ scripts/compare_acc_master.sh
The nek_gpu1 problem is setup for 256 elements and a polynomial order of 16. You may alter this setup in the SIZEanddata.reafiles as described inUSERGUIDE.pdf`.
$ cd test/nek_gpu1 $ cp ../../bin/makenek.mpi.titan . $ cp ../../bin/nekpmpi.titan . $ ./makenek.mpi.titan $ ./nekpmpi.titan data 16 1
$ cd test/nek_gpu1 $ cp ../../bin/makenek.acc.titan . $ cp ../../bin/nekpgpu.titan . $ ./makenek.acc.titan $ ./nekpgpu.titan data 1 1
$ cd test/nek_gpu1 $ cp ../../bin/makenek.cuda.titan . $ cp ../../bin/nekpgpu.titan . $ ./makenek.cuda.titan $ ./nekpgpu.titan data 1 1
[1] Jing Gong, Stefano Markidis, Erwin Laure, Matthew Otten, Paul Fischer, Misun Min,
Nekbone performance on GPUs with OpenACC and CUDA Fortran implementations,
The jounral of Supercomputing, Vol. 72, pp. 4160-4180, 2016.
[2] Matthew Otten, Jing Gong, Azamat Mametjanov, Aaron Vose, John Levesque, Paul
Fischer, and Misun Min, An MPI/OpenACC implementation of a high order
electromagnetics solver with GPUDirect communication, The International Journal
of High Performance Computing Application, Vol. 30, No. 3, pp. 320-334, 2016.
| Back | FazBrowse Home | New Git URL |