Traineeship

HPC-enabled AI Workflows for Drug Discovery and Computational Biology

Traineeship details

About the hosting organisation

Host institute / company

SoftMining SRL

Country

Sector(s)

Industry — SME (Small / Medium Enterprise)

Duration and availability

Duration

3 months

Earliest start date

01/03/2027

About the traineeship

Traineeship name

HPC-enabled AI Workflows for Drug Discovery and Computational Biology

Proposal ID

HPCTRAIN-OFFER-202601-029

Thematic area(s)

AI & Machine Learning, Bioinformatics & Computational Biology, Healthcare & Medical Imaging

Name of the position

HPC-enabled AI Trainee for Computational Drug Discovery

About the position

The trainee will join SoftMining’s technical and scientific team as an HPC-enabled AI trainee, working on computational workflows for the ADAPT platform. The trainee will support the design, execution, benchmarking, and documentation of scalable pipelines for AI-assisted drug discovery and computational biology. The role will combine software development, data analysis, workflow execution, and performance-aware implementation. The trainee will contribute to activities such as preparing and processing biomedical or molecular datasets, configuring computational experiments, running and monitoring batch workflows, evaluating pipeline performance, supporting AI-driven analysis, and documenting reproducible procedures. The trainee will work under the supervision of an internal mentor and will interact with team members involved in AI, computational biology, software development, and drug discovery activities. The position is intended to provide hands-on experience in how AI and data-driven methods are translated into professional HPC-enabled workflows for real-world biomedical and pharmaceutical R&D applications.

Minimum qualification level

Master's student, PhD candidate, The ideal candidate should have a background in Computer Science, Data Science, Bioinformatics, Computational Biology, Biomedical Engineering, Chemistry, Pharmaceutical Sciences, Physics, Mathematics, or a related technical/scientific field. Required or preferred skills include good programming skills, preferably in Python; basic knowledge of machine learning, data analysis, or scientific computing; familiarity with Linux command-line environments; and interest in HPC, cloud computing, distributed workflows, or scalable data-intensive applications. Previous experience with Jupyter Notebooks, Git, scientific Python libraries, APIs, databases, or workflow scripting is considered an advantage. Basic knowledge of biology, chemistry, omics data, molecular modelling, or drug discovery is useful but not mandatory. The candidate should be able to work in English and be motivated to operate in a multidisciplinary environment involving AI, software engineering, computational biology, and pharmaceutical R&D.

Learning objectives

Code porting & benchmarking, Data analysis & visualisation at scale, Scientific software engineering, The trainee will also gain practical exposure to the translation of AI and data-driven research methods into reproducible computational workflows for drug discovery and computational biology. In particular, the trainee will learn how to organise scientific data pipelines, document computational experiments, manage input/output data, and ensure traceability of results in a professional R&D environment. The traineeship will also develop the trainee’s ability to work across disciplines, connecting computer science, AI, bioinformatics, and pharmaceutical innovation. Additional learning outcomes include understanding the practical constraints of deploying AI methods on HPC or cloud-HPC environments, improving workflow reproducibility, identifying computational bottlenecks, and communicating technical results through clear documentation and internal reporting.

Tools and technologies to be used

Git / GitLab, Linux / Bash, Python, PyTorch, Python scientific stack, including pandas, NumPy, scikit-learn, and Jupyter Notebooks; Git-based version control; Linux command-line tools; REST APIs and database access tools; workflow scripting for data preprocessing, batch execution, and post-processing; Docker or similar container technologies for reproducible environments; RDKit and molecular data processing tools where applicable; biomedical and chemical databases such as ChEMBL, PDB, and literature-derived resources; internal ADAPT platform components for AI-assisted drug discovery and computational biology. Depending on the specific activity and available infrastructure, the trainee may also be exposed to cloud-HPC resources, workload managers such as Slurm, batch processing environments, workflow monitoring tools, and benchmarking utilities for evaluating runtime, scalability, and reproducibility.

View the HPCTRAIN call for trainees call text

Description

The trainee will work on HPC-enabled AI workflows supporting SoftMining’s ADAPT platform for computational drug discovery and computational biology. The work will focus on practical, hands-on activities related to the preparation, execution, optimisation, benchmarking, and documentation of scalable computational pipelines. The trainee will contribute to the preparation and preprocessing of molecular, biomedical, biological, and literature-derived datasets; the configuration and execution of AI-assisted analysis workflows; the support of large-scale virtual screening or candidate prioritisation activities; and the post-processing and interpretation of computational outputs. The trainee may also develop scripts or notebooks to automate data processing, compare alternative workflow configurations, and support reproducible reporting. The HPC dimension is central to the traineeship because the activities involve data-intensive and computationally demanding tasks that require scalable execution, batch processing, efficient resource use, and reproducible workflows. The trainee will learn how AI and machine learning methods can be integrated into HPC or cloud-HPC environments to process large input spaces, optimise computational experiments, and support decision-making in drug discovery and bioinformatics. The expected outcome is a documented and reproducible workflow component, benchmark, or technical report showing how selected AI- driven computational biology or drug discovery tasks can be executed, monitored, and improved in an HPC-enabled setting.

Submit your proposal

Please find more details on how to apply below.

Go to application portal