Published October 6, 2026 | Version v1.0.0-2025

Nextflow for HPC trainers materials

  • 1. Sydney Informatics Hub, University of Sydney

Description

This training material is designed to help other trainers re-run and adapt the 'Nextflow for HPC' workshop for their own learners. It includes ready-to-use lessons and guidance for delivering a two-part workshop on running reproducible and scalable scientific workflows with Nextflow on high performance computing (HPC) systems.

Nextflow is a popular bioinformatics workflow orchestrator that supports portable, reproducible, and scalable analysis across different computational infrastructures. It integrates with HPC job schedulers such as PBS Pro and Slurm, manages parallel task execution, and supports software containers for reproducible analysis.

The workshop materials are structured in two parts, using whole genome sequencing (WGS) short variant calling as a running example. Part one covers conceptual HPC foundations (shared systems, schedulers, software modules and containers, resource requests and job costs) and applies them to configuring and running the nf-core/sarek pipeline on HPC. Part two provides hands-on practice configuring, profiling, optimising, and scaling a custom multi-sample Nextflow workflow on HPC. Example code is provided for NCI Gadi (PBS Pro) and Pawsey Setonix (Slurm) and can be adapted to other HPC systems.

Format: Originally delivered as two half-day sessions. Adaptable for online, in-person, distributed/hybrid, or self-paced delivery.

Target audience: Trainers delivering Nextflow workshops, facilitators supporting hands-on training events, and contributors developing or maintaining Nextflow training curricula. The workshop itself is aimed at researchers and bioinformaticians who have run simple Nextflow pipelines and want to scale them up on HPC.

Prerequisites: Completion of an introductory Nextflow workshop (e.g. Hello Nextflow) or equivalent experience, basic command-line skills, and familiarity with how HPC clusters work (e.g. job submission, compute nodes).

Learning outcomes: By the end of the workshop, learners should be able to describe what an HPC is and recognise when a workflow needs one; load software with environment modules and run tools in Singularity containers; submit and monitor scheduler jobs and inspect their resource usage; explain how CPU, memory, and walltime requests affect scheduling and cost; distinguish multi-threading from scatter-gather parallelisation; configure Nextflow executors, queues, and containers to run workflows on HPC; layer institutional and run-level configuration files to tune nf-core and custom pipelines; use Nextflow reports, timelines, and trace files to profile resource usage; implement multi-threading and scatter-gather in a Nextflow workflow; and scale an optimised workflow to multiple samples.

Contact: georgina.samaha@sydney.edu.au

Notes

Technical requirements: participants require a personal computer with VSCode (with the Remote - SSH extension) or a comparable code editor, an SSH client, a web browser, and an account on the training HPC system (the materials use NCI Gadi and Pawsey Setonix). The training HPC must provide Nextflow, Singularity/Apptainer, and a supported job scheduler (e.g. PBS Pro or Slurm), with sufficient compute allocation and scratch storage for all participants. Before the workshop, participants should install required software/extensions and confirm they can connect to their assigned HPC.

Files

Sydney-Informatics-Hub/nf4hpc-materials-v1.0.0-2025.zip

Files (11.4 MB)

Additional details