Published August 31, 2012 | Version v1

D12.4: Performance Optimized Lustre

  • 1. BSC

Description

This report documents the research and development carried on within Task 12.4. The main goal of the task is to identify and address some open issues in file systems for multi-petascale and exascale facilities, aiming to the development of solutions that can be applied to the Lustre file system.
The addressed issues can be classified into two main areas: metadata management and data management. Metadata handling involves dealing with huge numbers of files and their hierarchical organization according the user’s view (including directory management and file attributes). Data handling deals with the storage of file contents and management data; this includes, in particular, techniques for automatic (self-tuned) placement of data on a system with many heterogeneous devices, aiming at maximizing bandwidth and minimizing response time.
The work carried on in the area of metadata management included the observation, measurement and study of a large scale system currently in production, in order to identify the key metadata-related issues; the development of a prototype aimed to improve the metadata behaviour in such system and also to provide a framework to easily deploy novel metadata management techniques on top of other systems; the measurement and study of specially deployed Lustre and GPFS prototypes to validate the presence of the metadata issues observed in current in-production systems; and finally the porting of the framework prototype to test novel metadata management techniques on the Lustre prototype facility.
In this line we have observed that in both Lustre and GPFS there are some scalability issues that reduce the performance of metadata operation when many files are used by the applications or when the number of accessing clients grows. The most important observation is that the number of files needed for the problem to appear is only a few hundreds and the number of clients a few dozens. This clearly shows that the problem needs to be addressed.
After our mechanism has been added to the GPFS system, we have observed that the decrease in performance that appears in the evaluated cases disappears making the system much more scalable with the number of files and clients. The main reason for this beneficial effect is that our middleware is able to convert not optimized cases (from GPFS point of view) into the optimized cases of GPFS.
Regarding data management, the work carried on in the present Task is based in the study of the limitations of current data distribution strategies, especially when dealing with sustained growth of storage capacity.
The results consist of a proposal for a novel data re-distribution technique for increased storage capacity, aimed to maximize bandwidth and responsiveness while minimizing the cost of data re-distribution. This technique that redistributes data using an approach that mixes the two current trends (deterministic and randomized placement) is able to achieve a perfect redistribution of data with minimal data movement (like the randomized approaches) and without the negative effects in metadata size or computing time of traditional randomized approaches.
It is also a contribution of this work a new level in the caching hierarchy that uses a tiny portion of all available storage systems to cache data and achieves performance results similar to the ones obtained when the data is perfectly well distributed, but starting from data very badly distributed.

Files

2IP-D12.4.pdf

Files (3.2 MB)

Name Size Download all
md5:0f1bb11d79ae3d6fb4e1c7380518e170
3.2 MB Preview Download

Additional details

Funding

European Commission
PRACE-2IP - PRACE - Second Implementation Phase Project 283493