Traineeship

HPC Trainee – GPU Optimisation and Scientific Workflow Integration

Traineeship details

About the hosting organisation

Host institute / company

LECAD Laboratory, Faculty of Mechanical Eng., University of Ljubljana

Country

Sector(s)

Academic / Higher Education

Duration and availability

Duration

6 months

Earliest start date

01/09/2026

About the traineeship

Traineeship name

HPC Trainee – GPU Optimisation and Scientific Workflow Integration

Proposal ID

HPCTRAIN-OFFER-202601-022

Thematic area(s)

Energy & Sustainability Modelling, Scientific Visualisation, High-Performance Computing (HPC) with specific focus on GPU accelerators and Scientific Workflow Optimisation. The trainee will attempt to accelerate different simulation steps of the workflow using GPU-accelerators and will also attempt to optimize connectivity between multiple simulation steps via efficient data passing. HPC is the central focus: all work will be carried out on multi-node GPU clusters using industry-standard parallel programming models.

Name of the position

HPC Trainee – GPU Optimisation and Scientific Workflow Integration

About the position

The trainee will join an active research/engineering team in the field of nuclear fusion and vast HPC experience. He will take on the role of an HPC-related software developer. He will contribute to two tightly linked objectives: 1. GPU optimisation and parallelisation – profiling existing simulation codes, identifying code bottlenecks, and implementing GPU kernels or accelerated routines to improve time-to-solution on HPC clusters. 2. Workflow integration – designing and implementing efficient, robust data-passing mechanisms between separate simulation codes (e.g., via shared memory, parallel I/O, coupling libraries, or in-situ data transfer), so that the outputs of one code can be consumed by the next with minimal latency and I/O overhead. The focus will be on three different simulation steps that work sequentially and pass data from one step to another: (1) field-line tracing algorithms for assessment of plasma power deposition to first wall in magnetically confined nuclear devices. Data from this step is passed to (2) thermal modelling code to estimate temperature distributions to the first wall. Finally, (3) ray-tracing algorithms simulating IR camera response will deliver final synthetic camera image.

Minimum qualification level

Bachelor's degree completed, Master's student, PhD candidate, PhD awarded, Postdoctoral / early-career researcher, Industry professional (no academic minimum), Field of study: Computer Science, Computational Science / Engineering, Physics, Mathematics, or similar. Essential skills: - Programming in C/C++; Python scripting - Basic understanding of parallel programming (MPI, OpenMP) - Familiarity with Linux / HPC cluster environments Desirable skills: - Experience with GPU programming (CUDA, HIP, OpenCL, or cross-platform libraries, such as ALPAKA, StarPU,...) - Knowledge of parallel I/O (HDF5, NetCDF, MPI-IO) or data-coupling libraries (e.g., ADIOS2, preCICE, MUSCLE3) - Exposure to profiling and performance-analysis tools is a bonus(e.g., Nsight, Score-P, TAU, VTune)

Learning objectives

Benchmark HPC simulation codes on GPUs and identify performance bottlenecks, Implement GPU code to accelerate computations, Parallel programming skills (MPI + GPU, OpenMP offloading) to scale simulation codes, Implement an efficient data-exchange code between two or more coupled simulation steps, Use HPC workflow and tools (Slurm, profilers, monitoring dashboards), Document technical work to a professional standard and present results to the team

Tools and technologies to be used

C / C++, CMake, CUDA, Git / GitLab, MPI, OpenACC, OpenMP, Python, Slurm, Parallel I/O & coupling: HDF5, ADIOS2 (or equivalent), MPI-IO

View the HPCTRAIN call for trainees call text

Description

The trainee will work on two main tasks: A) GPU optimisation and parallelisation The trainee will design GPU-accelerated code to replace bottlenecks in existing simulation codes. This includes writing and tuning CUDA/HIP kernels, exploiting GPU-aware MPI for inter-node communication, and benchmarking achieved versus roofline performance. The goal is a measurable reduction in wall-clock time for representative production runs. B) Multi-code workflow integration Simulation steps pass data between distinct codes. The trainee will analyse the current I/O-code, identify latency and memory bottlenecks, and implement an optimised in-memory or in-situ coupling layer. Solutions may involve shared-memory segments, ADIOS2 staging, or a lightweight coupling library, and will be validated for correctness and scalability.The exact techniques and tools to be used will be discussed and decided together with the team of experts in the laboratory.

Submit your proposal

Please find more details on how to apply below.

Go to application portal