{"id":226,"date":"2024-11-16T08:00:09","date_gmt":"2024-11-16T08:00:09","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=226"},"modified":"2024-11-17T00:43:26","modified_gmt":"2024-11-17T00:43:26","slug":"2018-19-ropenmp-a-resilient-parallel-programming-model-for-heterogeneous-systems","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=226","title":{"rendered":"2018-19: rOpenMP: A Resilient Parallel Programming Model for Heterogeneous Systems"},"content":{"rendered":"<p>Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale supercomputing. Extreme heterogeneity, i.e., using multiple, and potentially configurable, types of processors, accelerators and memory\/storage in a single computing platform, is adding a significant amount of complexity to the supercomputer hardware\/software ecosystem. Errors and failures reported by such heterogeneous hardware will need to be handled by the appropriate software component to enable efficient masking, recovery, and avoidance with little burden on the user.<\/p>\n<p>This project takes a first step toward resilience in leadership-class supercomputers with extreme heterogeneity. It performs research to enable fine-grain resilience for graphics processing units (GPU) accelerated systems, such as Oak Ridge National Laboratory&#8217;s Summit supercomputer, that is more efficient than traditional application-level checkpoint\/restart. The approach centers on a novel concept for Quality of Service (QoS) and corresponding extensions for the for OpenMP parallel programming model. This project develops (1) error and failure models, (2) software resilience strategies and protection domains, (3) OpenMP QoS language extensions for resilience, (4) OpenMP QoS runtime extensions and policies for resilience, and (5) a proof-of-concept prototype demonstrating these capabilities on the Summit supercomputer at Oak Ridge National Laboratory.<\/p>\n<p>The ultimate goal is to make fault resilience an integral part of the supercomputer hardware\/software ecosystem, such that the burden for providing it is on the system by design and not on the user as an afterthought.<\/p>\n<p align=\"center\"><img decoding=\"async\" src=\"images\/ropenmp\/interactions.png\" hspace=\"0\" vspace=\"0\" height=\"50%\" width=\"50%\"><br \/>\n<i>Figure: Compile-time workflow and run-time interactions of the rOpenMP prototype using LLVM 7<\/i><\/p>\n<h4>Funding Sources<\/h4>\n<ul>\n<li>Laboratory Directed Research and Development, <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\">Oak Ridge National Laboratory<\/a>\n<\/li>\n<\/ul>\n<h4>Participants<\/h4>\n<ul>\n<li>Christian Engelmann (PI), Geoffroy Vall\u00e9e, and Swaroop Pophale &#8212; <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\">Oak Ridge National Laboratory<\/a>\n<\/li>\n<\/ul>\n<h4>Peer-reviewed Workshop Publications<\/h4>\n<ol>\n<li>Christian Engelmann, Geoffroy R. Vall&eacute;e, and Swaroop Pophale. <b>Concepts for OpenMP Target Offload Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/parallel.auckland.ac.nz\/iwomp2019\" target=\"parallel.auckland.ac.nz\/iwomp2019\">15th International Workshop on OpenMP (IWOMP) 2019<\/a><\/i>, pages 78-93, Auckland, New Zealand, September 11-13, 2019. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-030-28595-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-28596-8_6\" target=\"publication\">10.1007\/978-3-030-28596-8_6<\/a>. <a href=\"javascript:showAbstract('Recent reliability issues with one of the fastest supercomputers in the world, Titan at Oak Ridge National Laboratory, demonstrated the need for resilience in large-scale heterogeneous computing. OpenMP currently does not address error and failure behavior. This paper takes a first step toward resilience for heterogeneous systems by providing the concepts for resilient OpenMP offload to devices. Using real-world error and failure observations, the paper describes the concepts and terminology for resilient OpenMP target offload, including error and failure classes and resilience strategies. It details the experienced general-purpose computing on graphics processing units errors and failures in Titan. It further proposes improvements in OpenMP, including a preliminary prototype design, to support resilient offload to devices for efficient handling of errors and failures in heterogeneous high-performance computing systems');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19concepts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann19concepts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19concepts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Talks and Lectures<\/h4>\n<ol>\n<li>Christian Engelmann. <b>Resilience in Parallel Programming Environments<\/b>. Invited talk at the <a href=\"http:\/\/iadac.github.io\/events\/adac8\" target=\"iadac.github.io\/events\/adac8\">8th Accelerated Data Analytics and Computing (ADAC) Institute Workshop<\/a>, Tokyo, Japan, October 30-31, 2019. <a href=\"javascript:showAbstract('Recent reliability issues with one of the fastest supercomputers in the world, Titan at Oak Ridge National Laboratory, demonstrated the need for resilience in large-scale heterogeneous computing. OpenMP currently does not address error and failure behavior. The presented work takes a first step toward resilience for heterogeneous systems by providing the concepts for resilient OpenMP offload to devices. Using real-world error and failure observations, this work describes the concepts and terminology for resilient OpenMP target offload, including error and failure classes and resilience strategies. It details the experienced general-purpose computing on graphics processing units errors and failures in Titan. It further proposes improvements in OpenMP, including a preliminary prototype design, to support resilient offload to devices for efficient handling of errors and failures in heterogeneous high-performance computing systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19resilience3.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19resilience3\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale supercomputing. Extreme heterogeneity, i.e., using multiple, and potentially configurable, types of processors, accelerators and memory\/storage in a single computing platform, is adding a significant amount of complexity to the supercomputer hardware\/software ecosystem. Errors and failures&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":145,"menu_order":19,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-226","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/226","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=226"}],"version-history":[{"count":10,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/226\/revisions"}],"predecessor-version":[{"id":1214,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/226\/revisions\/1214"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/145"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=226"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}