Traineeship
RISE Research Institutes of Sweden
Research Institute / National Laboratory
6 months
01/10/2026
HPC-AI Trainee: Development, Performance Engineering and Optimisation of Large-Scale AI Applications
HPCTRAIN-OFFER-202601-127 4/
AI & Machine Learning, High-Performance Data Analytics, Parallel Programming & Code Optimisation, Scientific Visualisation, Development, GPU Computing, Performance Engineering, Benchmarking, Performance Optimisation, Distributed Training, Scalability Analysis, Parallel Computing, AI Workflows, Performance Portability, Scientific Computing, HPC Performance Analysis, Supercomputing Architectures & software stack, AMD GPUs, NVIDIA GPUs, CPU-GPU porting and tunning.
HPC-AI Trainee
The trainee will join the High-Performance Computing (HPC) Group at RISE Research Institutes of Sweden and work alongside HPC and AI experts on the development, optimisation, benchmarking and performance analysis of Artificial Intelligence applications on modern heterogeneous EuroHPC systems. The traineeship will focus on understanding how AI workloads perform across different HPC architectures, including CPU-based and GPUaccelerated systems (AMD, NVIDIA) available through European HPC infrastructures. The trainee will gain hands-on experience in deploying, profiling, optimising and benchmarking AI applications on multiple hardware platforms while investigating scalability, portability and performance-efficiency trade-offs. Working with real-world scientific and industrial AI workloads, the trainee will evaluate how architectural differences influence application behaviour and will develop optimisation strategies that improve performance, resource utilisation and energy efficiency. The work will contribute to the development of best practices for running AI applications on next-generation European supercomputing infrastructures and support the broader adoption of HPC technologies for AI-driven research and innovation.
Bachelor's degree completed, Master's student, PhD candidate, PhD awarded, Postdoctoral / early-career researcher, Industry professional (no academic minimum), The ideal candidate is a master student, PhD Student, recent graduate, or early-career professional with a background in Computer Science, Data Science, Artificial Intelligence, Computational Science, Software Engineering, Applied Mathematics, or Computer Engineering. Candidates should have programming experience in Python, basic Linux command-line skills, an understanding of machine learning and deep learning concepts, and a strong interest in High-Performance Computing (HPC) and large-scale computing systems. Experience with AI frameworks such as PyTorch and TensorFlow is desirable, along with knowledge of parallel computing concepts, familiarity with GPU and accelerator technologies, and experience working in Linux environments. Knowledge of Git, software development workflows, and container technologies is also considered an advantage.
Code porting & benchmarking, Data analysis & visualisation at scale, GPU & accelerator programming, Industry collaboration & soft skills, Machine learning workflows on HPC, Performance optimisation & profiling, Research methodology in HPC contexts
CUDA, Linux / Bash, MPI, Python, PyTorch, Singularity / Apptainer, Slurm, Spack / EasyBuild, TensorFlow, The trainee will use programming languages and AI frameworks including Python, PyTorch, TensorFlow, and optimising AI applications. The traineeship will provide hands-on experience with HPC technologies such as Slurm, MPI, OpenMP, Apptainer/Singularity, and Docker for managing and executing workloads on large-scale computing systems. For performance analysis and optimisation, the trainee will work with tools including ROCm, ROCProfiler, Nsight Systems (where applicable), Grafana dashboards, and benchmarking frameworks. Software development will be supported through Git, GitLab/GitHub, and Linux-based development environments. The trainee will conduct experiments and benchmarking on AMD and NVIDIA GPU-based systems, as well as HPC clusters available through RISE collaborations and European HPC infrastructures.
View the HPCTRAIN call for trainees call textThe trainee will develop practical skills in High-Performance Computing, AI performance engineering and heterogeneous computing environments. By the end of the traineeship, the trainee will understand how to deploy and optimise AI applications on different HPC architectures and systematically evaluate their performance characteristics. The trainee will learn advanced profiling and benchmarking techniques, identify computational bottlenecks, optimise resource utilisation and conduct scalability studies across multiple computing platforms. A particular focus will be placed on understanding the impact of architectural differences in CPU, GPU - AMD / NVIDIA systems and their software stack and how these differences affect AI training and inference workloads. The traineeship will also provide experience with distributed computing, performance portability, reproducible benchmarking methodologies and scientific software development. The trainee will gain exposure to state-of-the-art European supercomputing infrastructures and acquire the skills required to support AI applications in modern HPC environments. There is also an opportunity for trainees to work on Sweden's HPC system, Arrhenius, in collaboration with the National Academic Infrastructure for Supercomputing in Sweden (NAISS), to carry out the proposed work during the traineeship.