Traineeship

HPC administration and software engineering

Traineeship details

About the hosting organisation

Host institute / company

LECAD, Faculty of Mechanical Eng

Country

Sector(s)

Academic / Higher Education

Duration and availability

Duration

6 months

Earliest start date

01/09/2026

About the traineeship

Traineeship name

HPC administration and software engineering

Proposal ID

HPCTRAIN-OFFER-202601-023

Thematic area(s)

HPC Systems Administration & Operations, High-Performance Computing (HPC) proficiency with specific focus on research software engineering approaches from ground up. The training covers HPC administration from node provisioning with Confluent, system software deployment with Ansible scripts and Steampunk checker, building software modules with Easybuild and Spack, building of large portable software applications with Apptainer for Appimage.

Name of the position

HPC orchestrator

About the position

The trainee will join an active research/engineering team in the field of nuclear fusion and vast HPC experience. He/She will take on the role of an HPC-related administration tasks that consist of understanding HPC cluster and components from filesystems, nodes , monitoring and control through job scheduler. Two main objectives are envisaged for HPC professional position. Firstly, administration at the hardware level with software tools for node-level engineering (xCat, SLURM, Ansible) will be trained and secondly, software stacks builds will be developed for HPC users and developers with ubiquitously used EasyBuild and Spack tools that will be packed with containers (Apptainer, Appimage) for building complex portable apps.

Minimum qualification level

Bachelor's degree completed, Master's student, PhD candidate, PhD awarded, Postdoctoral / early-career researcher, Industry professional (no academic minimum), Field of study: Computer Science, Computational Science / Engineering, Physics, Mathematics, or similar. Level: Master's student or PhD students. Essential skills: - Linux and Python scripting, Basic understanding of parallel programming (MPI, OpenMP) - Familiarity with Linux / HPC cluster environments Desirable skills: Experience with GPU paradigms (CUDA, HIP), YAML, Make, Git, Knowledge of parallel I/O, YAML

Learning objectives

Domain-specific application development, HPC system administration, Scientific software engineering, Workflow & job management, Node provisioning and node monitoring techniques, Understand cluster maintenance with SLURM and user/job management, Develop new and adjust software modules, Develop containers for large deployments with continuous integration, Use HPC workflow and tools (SLURM, profilers, monitoring dashboards)

Tools and technologies to be used

CMake, CUDA, Git / GitLab, HIP / ROCm, MPI, OpenMP, Python, Slurm, Spack / EasyBuild, Programming languages: YAML, Bash, Ansible, Python, Parallel I/O & coupling: HDF5, MPI-IO, HPC environment: multi-node GPU cluster, SLURM job scheduler

View the HPCTRAIN call for trainees call text

Description

The trainee will work on two main tasks: A) HPC node-level provisioning and maintenance The trainee will adopt complete chain of HPC node deployment from stock Enterprise Linux image, to building of node-level drivers for IO and GPU with Ansible automation. This includes writing and tuning Ansible scripts for building CUDA/HIP kernels, exploiting GPU-aware MPI for inter-node communication, and Lustre with Infiniband. The goal is proficiency in node-level maintenance, monitoring and upgrades within SLURM scheduler. B) HPC cluster-level software engineering Providing prebuild software modules tailored for target cluster will be main topic here with toolchains on different versions using reproducible EasyBuild scripts or developer like scripts in Spack, both in Python-like configs. Larger software are difficult to migrate due to library version interdependences and patches required. For that HPC containers will be exercised for different Linux targets. The exact techniques and tools to be used will be discussed and decided together with the team of experts in the laboratory.

Submit your proposal

Please find more details on how to apply below.

Go to application portal