Conference paper Open Access

Optimum Interval for Application-level Checkpoints

Siavvas, Miltiadis; Gelenbe, Erol

Checkpointing is commonly adopted for enhancing the performance of software applications that operate in the presence of failures. Among the existing checkpointing strategies, Application-level Checkpoint and Restart (ALCR) is considered the most efficient, since it leaves smaller memory footprint, but it requires significant development effort. Although existing ALCR tools and libraries manage to reduce the effort required for implementing the checkpoints, they do not provide recommendations regarding their inter-checkpoint interval. To this end, in the present paper, we develop a mathematical model to estimate the optimum checkpoint interval, i.e., the interval between two successive checkpoints that minimises the average execution time of the application. The case of programs with loops and nested loops is also discussed. The results are illustrated with several numerical examples.

Files (454.6 kB)
Name Size
454.6 kB Download
All versions This version
Views 9797
Downloads 141141
Data volume 64.1 MB64.1 MB
Unique views 8888
Unique downloads 137137


Cite as