Abstract
Memory resources are an important aspect to consider when designing high performing programs. This is especially true for programs running on graphical processing units, GPUs, yet this is not something trivially done using current OpenMP target offloading. In this paper, we examine methods for implementing parallel programs running on GPUs, which rely on locally shared memory resources and intricate synchronization. Employing the methods, we show you can achieve between 1.5 to 9 relative speedup over a range of compilers. We evaluate portability by running experiments on two systems, utilizing different GPU technologies and vendors. We further investigate scheduling, synchronization and execution time of our experiments, to better understand the overhead associated with using OpenMP, compared to architecture specific languages. Lastly, we argue that improved GPU scheduling could yield a potential speedup of 3.
| Original language | English |
|---|---|
| Title of host publication | 19th International Workshop on OpenMP |
| Volume | 14114 |
| Publisher | Springer |
| Publication date | 2023 |
| Pages | 114-128 |
| ISBN (Print) | 978-3-031-40743-7 |
| ISBN (Electronic) | 978-3-031-40744-4 |
| DOIs | |
| Publication status | Published - 2023 |
| Event | 19th International Workshop on OpenMP - Bristol University, Bristol, United Kingdom Duration: 12 Sept 2023 → 15 Sept 2023 |
Workshop
| Workshop | 19th International Workshop on OpenMP |
|---|---|
| Location | Bristol University |
| Country/Territory | United Kingdom |
| City | Bristol |
| Period | 12/09/2023 → 15/09/2023 |
Keywords
- GPGPU Programming
- OpenMP Target Offloading
- Shared Memory
- Fine-Grained Parallelism
Fingerprint
Dive into the research topics of 'OpenMP Target Offload Utilizing GPU Shared Memory'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver