AGU Leptoukh Lecture - Data-driven science - the coherence of Big data, HPC, and informatics and crossing the next chasm
Description
The colocating of scientific datasets with HPC computational infrastructure has a long and steady timeline with deepening alignment that demonstrates success. Questions of whether to “bring data to the compute”, or “compute to the data” now has decades of maturity - considering and reconsidering the benefits, weaknesses and challenges both technically and socially. The historical disconnect in approaches between the sciences with large volume data and that of the long tail of data are also now steadily closing. The standards for interoperability and interconnectivity between scientific fields have been slowly maturing, and in many cases transdisciplinary science is now a reality. Furthermore, the emergence of hyper-scale cloud utility computing now makes the ability to colocate compute processing and data storage possible for many scientists.
As we approach the AGU Centenary, we clearly see the past looking different to the future. For example, when preparing for the exascale age it was known that there were going to be changes (hardware software, workflow and interconnectedness) that would require addressing many unknowns. There are now clearly more disruptions and major challenges to come for digital science and more broad applications for highly performant and resilient computing infrastructures.
The technical infrastructure for underpinning science is no longer advancing according to Moore’s law (and equivalents) and is evolving in unexpected ways. These have required us to consider our software development strategy, the plan for improvements and maintainability, and reconsidering some old assumptions of data precision and reproducibility. Furthermore, we are now seeing a glimpse of the non-von Neumann computing age, which may against test more assumptions to bypass existing bottlenecks.
The funding/business case and overall value proposition for celebrated open data and its FAIRness is being retested, despite the explosion in the interest and downstream opportunities that it represents. The needs of data-driven science with powerful technologies such as AI, deep learning, data analytics have much stronger requirements around the quality of data, information management and persistent exposure of its complexity. This is evolving at the same time as ubiquitous IOT, fog computing, and blockchain pipelines have emerged.
In this talk I will discuss the journey so far, and consider some of the chasms yet to cross.
Files
Files
(201.8 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:04c8a49205befbde865e9c6d709b4aa4
|
201.8 MB | Download |