ODISSEE - D5.1 : Preliminary Scaling Assessment of Users Code
Authors/Creators
Contributors
Project manager (2):
Project member (3):
Description
This document provides a high-level overview of the preliminary scalability assessment study for radio-astronomy workflows within the ODISSEE project, specifically focusing on preparations for the Square Kilometre Array Observatory (SKAO). The scalability of a reference radio-interferometry pipeline designed to process massive datasets by converting antenna "visibilities" into scientific images is examined. While the project's broader goal includes High Energy Physics (LHCb), this report focuses exclusively on radio-astronomy because LHCb workflows are "embarrassingly parallel" and do not require the same type of scalability study
The primary tools assessed include OSKAR (simulator), QuartiCal and KillMS (calibration), and DDFacet (imaging) and individual components of the overall workflow in the context of the work on composability (WP2).
Key Findings include:
-
Simulator Efficiency: OSKAR demonstrates near-ideal weak scaling on high-performance architectures like Jean Zay and the DALIA (NVIDIA GB200 NVL72) cluster. Optimization studies on DALIA showed that increasing the source chunk size for simulations can lead to a 2× speedup by reducing overhead and improving GPU utilization.
-
Pipeline Orchestration: Using pyCOMPSs for task-based management allows for the concurrent execution of simulation tasks on GPUs and calibration tasks on CPUs, effectively overlapping different stages of the workflow to maximize resource usage.
-
First demonstration of large scale multi-node deployment of a SelfCal and Imaging workflow (DDF-pipeline) on the state-of-the-art Jean Zay supercomputer showing good scaling properties with data volume, though deconvolution overhead increases as images get deeper.
-
Imaging components performance and scalability: GPU implementations of the gridding and degridding components achieved massive speedups of 40x to 138x compared to single-core CPU implementations. The NVIDIA H100 consistently delivered the best results, although performance is currently limited by memory bandwidth and atomic operations.
-
Calibration components performance and scalability: The study highlights a critical distinction in GPU programming models. "Cooperative" kernels (where multiple threads work on a single solve, such as in CohJones) significantly outperform "one-thread-per-solve" models. In some tests, cooperative designs turned a GPU performance deficit into a tenfold advantage.
Files
ODISSEE_D5.1 – Preliminary-Scaling-Assessment-VF (1).pdf
Files
(7.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:a6b2801accb3303cd7036aba2f220a38
|
7.7 MB | Preview Download |
Additional details
Funding
Dates
- Available
-
2026-07-08