{"id":373,"date":"2023-02-14T08:00:28","date_gmt":"2023-02-14T08:00:28","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=373"},"modified":"2023-02-15T20:13:13","modified_gmt":"2023-02-15T20:13:13","slug":"2002-04-super-scalable-algorithms-for-next-generation-high-performance-cellular-architectures","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=373","title":{"rendered":"2002-04: Super-Scalable Algorithms for Next-Generation High-Performance Cellular Architectures"},"content":{"rendered":"<p>This research in cellular architectures is part of a Cooperative Research and Development Agreement (CRADA) between IBM and Oak Ridge National Laboratory (ORNL) to develop algorithms for the next-generation of supercomputers. It focuses on the development of algorithms that are able to use a 100,000-processor machine efficiently and are capable of adapting to or simply surviving faults. Such huge computer systems, like the IBM Blue Gene\/L, need to address already existing problems in algorithm scalability and fault-tolerance, which continue to increase with system scale.<\/p>\n<p>In a first step, the team at ORNL develops a simulator to emulate up to 5,000 virtual processors on a single real processor, solving a simple equation at the virtual 5,000 processor scale. It a second step, the emulation is extended to 500,000 virtual processors on a cluster with 5 real processors, solving simple equations and performing advanced collective communication primitives.<\/p>\n<h4>Funding Sources<\/h4>\n<ul>\n<li>Laboratory Directed Research and Development, <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\" rel=\"noopener\">Oak Ridge National Laboratory<\/a>\n<\/li>\n<\/ul>\n<h4>Participating Institutions<\/h4>\n<ul>\n<li><a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\" rel=\"noopener\">Oak Ridge National Laboratory<\/a><\/li>\n<\/ul>\n<h4>Peer-reviewed Conference Publications<\/h4>\n<ol>\n<li>Christian Engelmann and George A. (Al) Geist. <b>Super-Scalable Algorithms for Computing on 100,000 Processors<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.iccs-meeting.org\/iccs2005\" target=\"www.iccs-meeting.org\/iccs2005\" rel=\"noopener\">5th International Conference on Computational Science (ICCS) 2005<\/a>, Part I<\/i>, pages 313-320, Atlanta, GA, USA, May 22-25, 2005. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\" rel=\"noopener\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-540-26032-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/11428831_39\" target=\"publication\" rel=\"noopener\">10.1007\/11428831_39<\/a>. Acceptance rate 35%. <a href=\"javascript:showAbstract('In the next five years, the number of processors in high-end systems for scientific computing is expected to rise to tens and even hundreds of thousands. For example, the IBM Blue Gene\/L can have up to 128,000 processors and the delivery of the first system is scheduled for 2005. Existing deficiencies in scalability and fault-tolerance of scientific applications need to be addressed soon. If the number of processors grows by a magnitude and efficiency drops by a magnitude, the overall effective computing performance stays the same. Furthermore, the mean time to interrupt of high-end computer systems decreases with scale and complexity. In a 100,000-processor system, failures may occur every couple of minutes and traditional checkpointing may no longer be feasible. With this paper, we summarize our recent research in super-scalable algorithms for computing on 100,000 processors. We introduce the algorithm properties of scale invariance and natural fault tolerance, and discuss how they can be applied to two different classes of algorithms. We also describe a super-scalable diskless checkpointing algorithm for problems that can`t be transformed into a super-scalable variant, or where other solutions are more efficient. Finally, a 100,000-processor simulator is presented as a platform for testing and experimentation.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05superscalable.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05superscalable.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05superscalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Peer-reviewed Workshop Publications<\/h4>\n<ol>\n<li>Christian Engelmann and George A. (Al) Geist. <b>A Diskless Checkpointing Algorithm for Super-scale Architectures Applied to the Fast Fourier Transform<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.cs.msstate.edu\/ clade2003\" target=\"www.cs.msstate.edu\/ clade2003\" rel=\"noopener\">Challenges of Large Applications in Distributed Environments Workshop (CLADE) 2003<\/a>, in conjunction with the <a href=\"http:\/\/csag.ucsd.edu\/HPDC-12\" target=\"csag.ucsd.edu\/HPDC-12\" rel=\"noopener\">12th IEEE International Symposium on High Performance Distributed Computing (HPDC) 2003<\/a><\/i>, pages 47, Seattle, WA, USA, June 21, 2003. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\" rel=\"noopener\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-1984-9. DOI <a href=\"http:\/\/dx.doi.org\/xpls\/abs_all.jsp?arnumber=4159902\" target=\"publication\" rel=\"noopener\">xpls\/abs_all.jsp?arnumber=4159902<\/a>. <a href=\"javascript:showAbstract('This paper discusses the issue of fault-tolerance in distributed computer systems with tens or hundreds of thousands of diskless processor units. Such systems, like the IBM Blue Gene\/L, are predicted to be deployed in the next five to ten years. Since a 100,000-processor system is going to be less reliable, scientific applications need to be able to recover from occurring failures more efficiently. In this paper, we adapt the present technique of diskless checkpointing to such huge distributed systems in order to equip existing scientific algorithms with super-scalable fault-tolerance. First, we discuss the method of diskless checkpointing, then we adapt this technique to super-scale architectures and finally we present results from an implementation of the Fast Fourier Transform that uses the adapted technique to achieve super-scale fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann03diskless.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann03diskless.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann03diskless\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Talks and Lectures<\/h4>\n<ol>\n<li>Christian Engelmann. <b>Advanced Fault Tolerance Solutions for High Performance Computing<\/b>. Seminar at the <a href=\"http:\/\/www.laas.fr\" target=\"www.laas.fr\" rel=\"noopener\">Laboratoire d&#39;Analyse et d&#8217;Architecture des Syst&eacute;mes<\/a>, <a href=\"http:\/\/www.cnrs.fr\" target=\"www.cnrs.fr\" rel=\"noopener\">Centre National de la Recherche Scientifique<\/a>, Toulouse, France, February 11, 2008. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08advanced.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08advanced\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Advanced Fault Tolerance Solutions for High Performance Computing<\/b>. Invited talk at the <a href=\"http:\/\/www.thaigrid.or.th\/wttc2007\" target=\"www.thaigrid.or.th\/wttc2007\" rel=\"noopener\">Workshop on Trends, Technologies and Collaborative Opportunities in High Performance and Grid Computing (WTTC) 2007<\/a>, Khon Kean, Thailand, June 8, 2007. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07advanced2.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07advanced2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Advanced Fault Tolerance Solutions for High Performance Computing<\/b>. Invited talk at the <a href=\"http:\/\/www.thaigrid.or.th\/wttc2007\" target=\"www.thaigrid.or.th\/wttc2007\" rel=\"noopener\">Workshop on Trends, Technologies and Collaborative Opportunities in High Performance and Grid Computing (WTTC) 2007<\/a>, Bangkok, Thailand, June 4-5, 2007. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07advanced.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07advanced\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Diskless Checkpointing on Super-scale Architectures &#8211; Applied to the Fast Fourier Transform<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/pp04\" target=\"www.siam.org\/meetings\/pp04\" rel=\"noopener\">11th SIAM Conference on Parallel Processing for Scientific Computing (SIAM PP) 2004<\/a>, San Francisco, CA, USA, February 25, 2004. <a href=\"javascript:showAbstract('This talk discusses the issue of fault-tolerance in distributed computer systems with tens or hundreds of thousands of diskless processor units. Such systems, like the IBM Blue Gene\/L, are predicted to be deployed in the next five to ten years. Since a 100,000-processor system is going to be less reliable, scientific applications need to be able to recover from occurring failures more efficiently. In this paper, we adapt the present technique of diskless checkpointing to such huge distributed systems in order to equip existing scientific algorithms with super-scalable fault-tolerance. First, we discuss the method of diskless checkpointing, then we adapt this technique to super-scale architectures and finally we present results from an implementation of the Fast Fourier Transform that uses the adapted technique to achieve super-scale fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann04diskless.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann04diskless\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Super-scalable Algorithms &#8211; Next Generation Supercomputing on 100,000 and more Processors<\/b>. Seminar at the <a href=\"http:\/\/www.csm.ornl.gov\" target=\"www.csm.ornl.gov\" rel=\"noopener\">Computer Science and Mathematics Division<\/a>, <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\" rel=\"noopener\">Oak Ridge National Laboratory<\/a>, Oak Ridge, TN, USA, January 29, 2004. <a href=\"javascript:showAbstract('This talk discusses recent research into the issues and potential problems of algorithm scalability and fault-tolerance on next-generation high-performance computer systems with tens and even hundreds of thousands of processors. Such massively parallel computers, like the IBM Blue Gene\/L, are going to be deployed in the next five to ten years and existing deficiencies in scalability and fault-tolerance need to be addressed soon. Scientific algorithms have shown poor scalability on 10,000-processor systems that exist today. Furthermore, future systems will be less reliable due to the large number of components. Super-scalable algorithms, which have the properties of scale invariance and natural fault-tolerance, are able to get the correct answer despite multiple task failures and without checkpointing. We will show that such algorithms exist for a wide variety of problems, such as finite difference, finite element, multigrid and global maximum. Despite these findings, traditional algorithms may still be preferred due to their known behavior, or simply because a super-scalable algorithm does not exist or is hard to find for a particular problem. In this case, we propose a peer-to-peer diskless checkpointing algorithm that can provide scale invariant fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann04superscalable.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann04superscalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>This research in cellular architectures is part of a Cooperative Research and Development Agreement (CRADA) between IBM and Oak Ridge National Laboratory (ORNL) to develop algorithms for the next-generation of supercomputers. It focuses on the development of algorithms that are able to use a 100,000-processor machine efficiently and are capable of adapting to or simply&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":145,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-373","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/373","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=373"}],"version-history":[{"count":4,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/373\/revisions"}],"predecessor-version":[{"id":418,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/373\/revisions\/418"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/145"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=373"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}