{"id":161,"date":"2025-02-01T08:00:17","date_gmt":"2025-02-01T08:00:17","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=161"},"modified":"2025-02-01T23:38:12","modified_gmt":"2025-02-01T23:38:12","slug":"2015-2019-catalog-characterizing-faults-errors-and-failures-in-extreme-scale-systems","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=161","title":{"rendered":"2015-19: Catalog: Characterizing Faults, Errors, and Failures in Extreme-Scale Systems"},"content":{"rendered":"<p>US Department of Energy (DOE) leadership computing facilities are in the process of deploying extreme-scale high-performance computing (HPC) systems with the long-range goal of building exascale systems that perform more than a quintillion (a billion billion) operations per second. More powerful computers mean researchers can simulate biological, chemical, and other physical interactions with an unprecedented amount of realism. However, as HPC systems become more complex, system integrators, component manufacturers as well as computing facilities have to and are preparing for unique computing challenges. Of particular concern are occurrences of unfamiliar or more frequent faults in both hardware technologies and software applications that can lead to computational errors or system failures.<\/p>\n<p>This project helps DOE computing facilities protect extreme-scale systems by characterizing potential faults and creating models that predict their propagation and impact. The Collaboration of Oak Ridge, Argonne and Lawrence Livermore National Laboratories (CORAL) is a private\/public partnership that will stand up three extreme-scale systems. By monitoring hardware and software performance on current DOE systems and applying the data to fault analysis and vulnerability studies, this effort captures observed and inferred fault conditions and extrapolate this knowledge to CORAL and other extreme-scale systems.<\/p>\n<p>Using these analyses, the project team creates assessment tools, including a fault taxonomy and catalog as well as fault models, to provide computing facilities with a clear picture of the fault characteristics in DOE computing environments and inform technical and operational decisions to improve resilience. For more information, please visit <a href=\"https:\/\/ornlwiki.atlassian.net\/wiki\/display\/CFEFIES\" target=\"ornlwiki.atlassian.net_wiki_display_CFEFIES\">ornlwiki.atlassian.net\/wiki\/display\/CFEFIES<\/a>.<\/p>\n<p align=\"center\"><img decoding=\"async\" src=\"images\/catalog\/comparison.png\" hspace=\"20\" vspace=\"0\" height=\"30%\" width=\"30%\"><img decoding=\"async\" src=\"images\/catalog\/fraction.png\" hspace=\"20\" vspace=\"0\" height=\"52%\" width=\"52%\"><br \/>\n<i>Figure: Analysis results of 1.2 billion node hours of logs from the Jaguar, Titan, and Eos systems at the Oak Ridge Leadership Computing Facility (OLCF)<\/i><\/p>\n<h4>Prominent Solutions<\/h4>\n<ul>\n<li><a href=\"?page_id=522\">Characterization of Faults, Errors, and Failures in Extreme-Scale Systems<\/a><\/li>\n<\/ul>\n<h4>Funding Sources<\/h4>\n<ul>\n<li>Resilience for Extreme Scale Supercomputing Systems Program, <a href=\"http:\/\/science.energy.gov\/ascr\" target=\"science.energy.gov_ascr\">Office of Advanced Scientific Computing Research<\/a>, Office of Science, U.S. Department of Energy\n  <\/li>\n<\/ul>\n<h4>Participants<\/h4>\n<ul>\n<li>Christian Engelmann (PI), Byung Hoon (Hoony) Park, Devesh Tiwari, Saurabh Gupta, and Mohit Kumar &#8212; <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\">Oak Ridge National Laboratory<\/a>\n<\/li>\n<li>Martin Schulz and Ignacio Laguna &#8212; <a href=\"http:\/\/www.llnl.gov\" target=\"www.llnl.gov\">Lawrence Livermore National Laboratory<\/a>\n<\/li>\n<li>Marc Snir, Franck Cappello, and Rinku Gupta &#8212; <a href=\"http:\/\/www.anl.gov\" target=\"www.anl.gov\">Argonne National Laboratory<\/a>\n<\/li>\n<\/ul>\n<h4>In the News<\/h4>\n<p><b>2021-01-04:<\/b> HPCwire. <a href=\"https:\/\/www.hpcwire.com\/2021\/01\/04\/whats-new-in-hpc-research-gpu-lifetimes-the-square-kilometre-array-support-tickets-more\" target=\"www.hpcwire.com\/2021\/01\/04\/whats-new-in-hpc-research-gpu-lifetimes-the-square-kilometre-array-support-tickets-more\">What&#8217;s New in HPC Research: GPU Lifetimes, the Square Kilometre Array, Support Tickets &amp; More<\/a>.<br \/>\n<b>2018-11-19:<\/b> HPCwire. <a href=\"https:\/\/www.hpcwire.com\/2018\/11\/19\/whats-new-in-hpc-research-thrill-for-big-data-scaling-resilience-and-more\" target=\"www.hpcwire.com_2018_11_19_whats-new-in-hpc-research-thrill-for-big-data-scaling-resilience-and-more\">What&#8217;s New in HPC Research: Thrill for Big Data, Scaling Resilience and More<\/a>.<br \/>\n<b>2018-08-05:<\/b> inside HPC. <a href=\"https:\/\/insidehpc.com\/2018\/08\/characterizing-faults-errors-failures-extreme-scale-computing-systems\/\" target=\"insidehpc.com_2018_08_characterizing-faults-errors-failures-extreme-scale-computing-systems\">Characterizing Faults, Errors and Failures in Extreme-Scale Computing Systems<\/a>.\n<\/p>\n<h4>Peer-reviewed Journal Publications<\/h4>\n<ol>\n<li>Emmanuel Agullo, Mirco Altenbernd, Hartwig Anzt, Leonardo Bautista-Gomez, Tommaso Benacchio, Luca Bonaventura, Hans-Joachim Bungartz, Sanjay Chatterjee, Florina M. Ciorba, Nathan DeBardeleben, Daniel Drzisga, Sebastian Eibl, Christian Engelmann, Wilfried N. Gansterer, Luc Giraud, Dominik G&ouml;ddeke, Marco Heisig, Fabienne J&eacute;z&eacute;quel, Nils Kohl, Xiaoye Sherry Li, Romain Lion, Miriam Mehl, Paul Mycek, Michael Obersteiner, Enrique S. Quintana-Ort&iacute;, Francesco Rizzi, Ulrich R&uuml;de, Martin Schulz, Fred Fung, Robert Speck, Linda Stals, Keita Teranishi, Samuel Thibault, Dominik Th&ouml;nnes, Andreas Wagner, and Barbara Wohlmuth. <b>Resiliency in Numerical Algorithm Design for Extreme Scale Simulations<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 36, number 2, pages 251-285, March 1, 2022. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/10943420211055188\" target=\"publication\">10.1177\/10943420211055188<\/a>. <a href=\"javascript:showAbstract('This work is based on the seminar titled &amp;#39;Resiliency in Numerical Algorithm Design for Extreme Scale Simulations' held March 1-6, 2020 at Schloss Dagstuhl, that was attended by all the writers. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 hours on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 1023 floating- point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications, and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/agullo22resiliency.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#agullo22resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar, Saurabh Gupta, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, and Devesh Tiwari. <b>Study of Interconnect Errors, Network Congestion, and Applications Characteristics for Throttle Prediction on a Large Scale HPC System<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 153, pages 29-43, July 1, 2021. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2021.03.001\" target=\"publication\">10.1016\/j.jpdc.2021.03.001<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of peta ops. The performance of these systems depends on various components like processors, memory, and interconnect. Among  all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar21study.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar21study\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Marc Snir, Robert W. Wisniewski, Jacob A. Abraham, Sarita V. Adve, Saurabh Bagchi, Pavan Balaji, Jim Belak, Pradip Bose, Franck Cappello, Bill Carlson, Andrew A. Chien, Paul Coteus, Nathan A. Debardeleben, Pedro Diniz, Christian Engelmann, Mattan Erez, Saverio Fazzari, Al Geist, Rinku Gupta, Fred Johnson, Sriram Krishnamoorthy, Sven Leyffer, Dean Liberty, Subhasish Mitra, Todd Munson, Rob Schreiber, Jon Stearley, and Eric Van Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 28, number 2, pages 127-171, May 1, 2014. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/1094342014522573\" target=\"publication\">10.1177\/1094342014522573<\/a>. <a href=\"javascript:showAbstract('We present here a report produced by a workshop on  Addressing failures in exascale computing&amp;#39; held in Park City,  Utah, 4-11 August 2012. The charter of this workshop was to  establish a common taxonomy about resilience across all the  levels in a computing system, discuss existing knowledge on  resilience across the various hardware and software layers  of an exascale system, and build on those results, examining  potential solutions from both a hardware and software  perspective and focusing on a combined approach. The workshop brought together participants with expertise in  applications, system software, and hardware; they came from  industry, government, and academia, and their interests ranged  from theory to implementation. The combination allowed broad  and comprehensive discussions and led to this document, which  summarizes and builds on those discussions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/snir14addressing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir14addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Peer-reviewed Conference Publications<\/h4>\n<ol>\n<li>George Ostrouchov, Don Maxwell, Rizwan Ashraf, Christian Engelmann, Mallikarjun Shankar, and James Rogers. <b>GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc20.supercomputing.org\" target=\"sc20.supercomputing.org\">33rd IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2020<\/a><\/i>, pages 41:1-14, Atlanta, GA, USA, November 15-20, 2020. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 9781728199986. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SC41405.2020.00045\" target=\"publication\">10.1109\/SC41405.2020.00045<\/a>. Acceptance rate 25.1% (95\/378). <a href=\"javascript:showAbstract('The Cray XK7 Titan was the top supercomputer system in the world for a very long time and remained critically important throughout its nearly seven year life. It was also a very interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three very significant rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 operation years in the GPU lifetimes, which correspond to Titan&amp;#39;s 6 year long productive period after an initial break-in period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the system cooling architecture and job scheduling. In addition to describing some of the system history, the data collection, data cleaning, and our analysis of the data, we provide reliability recommendations for designing future state of the art supercomputing systems and their operation. We make the data and our analysis codes publicly available.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ostrouchov20gpu.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ostrouchov20gpu.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ostrouchov20gpu\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar, Saurabh Gupta, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, and Devesh Tiwari. <b>Understanding and Analyzing Interconnect Errors and Network Congestion on a Large Scale HPC System<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.dsn.org\" target=\"www.dsn.org\">48th IEEE\/IFIP International Conference on Dependable  Systems and Networks (DSN) 2018<\/a><\/i>, pages 107-114, Luxembourg City, Luxembourg, June 25-28, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-5596-2. ISSN 2158-3927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DSN.2018.00023\" target=\"publication\">10.1109\/DSN.2018.00023<\/a>. Acceptance rate 27.2% (62\/228). <a href=\"javascript:showAbstract('Today&amp;#39;s High Performance Computing (HPC) systems are capable of delivering performance in the order of petaflops due to the fast computing devices, network interconnect, and back-end storage systems. In particular, interconnect resilience and congestion resolution methods have a major impact on the overall interconnect and application performance. This is especially true for scientific applications running multiple processes on different compute nodes as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks state-of-practice experience reports that detail how different interconnect errors and congestion events occur on large-scale HPC systems. Therefore, in this paper, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors and congestion events. We also study the interaction between interconnect, errors, network congestion and application characteristics.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar18understanding.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar18understanding\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Bin Nie, Ji Xue, Saurabh Gupta, Tirthak Patel, Christian Engelmann, Evgenia Smirni, and Devesh Tiwari. <b>Machine Learning Models for GPU Error Prediction in a Large Scale HPC System<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.dsn.org\" target=\"www.dsn.org\">48th IEEE\/IFIP International Conference on Dependable  Systems and Networks (DSN) 2018<\/a><\/i>, pages 95-106, Luxembourg City, Luxembourg, June 25-28, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-5596-2. ISSN 2158-3927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DSN.2018.00022\" target=\"publication\">10.1109\/DSN.2018.00022<\/a>. Acceptance rate 27.2% (62\/228). <a href=\"javascript:showAbstract('Recently, GPUs have been widely deployed on large-scale HPC systems to provide powerful computational capability for scientific applications from various domains. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative. Therefore, in this paper, we firstly study the conditions that trigger GPU errors with six-month trace data collected from a large-scale operational HPC system. Then, we resort to machine learning techniques to predict the occurrence of GPU errors, by taking advantage of the temporal and spatial dependency of the collected data. As discussed in the evaluation section, the prediction framework is robust and accurate under different workloads.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/nie18machine.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#nie18machine\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Gupta, Tirthak Patel, Christian Engelmann, and Devesh Tiwari. <b>Failures in Large Scale Systems: Long-term Measurement, Analysis, and Implications<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc17.supercomputing.org\" target=\"sc17.supercomputing.org\">30th IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2017<\/a><\/i>, pages 44:1-44:12, Denver, CO, USA, November 12-17, 2017. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-5114-0. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3126908.3126937\" target=\"publication\">10.1145\/3126908.3126937<\/a>. Acceptance rate 18.7% (61\/327). <a href=\"javascript:showAbstract('Resilience is one of the key challenges in maintaining high efficiency of future extreme scale supercomputers. Unfortunately, field-data based reliability studies are far in between and not exhaustive. Most HPC researchers and system practitioners still rely on outdated studies to understand HPC reliability characteristics and plan for future HPC systems. While the complexity of managing system reliability has increased, the public knowledge sharing about lessons learned from HPC centers has not increased in the same proportion. To bridge this gap, in this work, we compare and contrast the reliability characteristics of multiple large-scale HPC production systems, and discuss new take-aways and con rm previous findings which continue to be valid.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/gupta17failures.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/gupta17failures.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#gupta17failures\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Bin Nie, Ji Xue, Saurabh Gupta, Christian Engelmann, Evgenia Smirni, and Devesh Tiwari. <b>Characterizing Temperature, Power, and Soft-Error Behaviors in Data Center Systems: Insights, Challenges, and Opportunities<\/b>. In <i>Proceedings of the <a href=\"http:\/\/mascots2017.cs.ucalgary.ca\" target=\"mascots2017.cs.ucalgary.ca\">25th IEEE International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS) 2017<\/a><\/i>, pages 22-31, Banff, AB, Canada, September 20-22, 2017. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-2764-8. ISSN 2375-0227. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/MASCOTS.2017.12\" target=\"publication\">10.1109\/MASCOTS.2017.12<\/a>. Acceptance rate 30.95% (26\/84). <a href=\"javascript:showAbstract('GPUs have become part of the mainstream high performance computing facilities that increasingly require more computational power to simulate physical phenomena quickly and accurately. However, GPU nodes also consume significantly more power than traditional CPU nodes, and high power consumption introduces new system operation challenges, including increased temperature, power\/cooling cost, and lower system reliability. This paper explores how power consumption and temperature characteristics affect reliability, provides insights into what are the implications of such understanding, and how to exploit these insights toward predicting GPU errors using neural networks.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/nie17characterizing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#nie17characterizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Kun Tang, Devesh Tiwari, Saurabh Gupta, Ping Huang, QiQi Lu, Christian Engelmann, and Xubin He. <b>Power-Capping Aware Checkpointing: On the Interplay Among Power-Capping, Temperature, Reliability, Performance, and Energy<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.dsn.org\" target=\"www.dsn.org\">46th IEEE\/IFIP International Conference on Dependable  Systems and Networks (DSN) 2016<\/a><\/i>, pages 311-322, Toulouse, France, June 28 &#8211; July 1, 2016. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISSN 2158-3927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DSN.2016.36\" target=\"publication\">10.1109\/DSN.2016.36<\/a>. Acceptance rate 22.4% (58\/259). <a href=\"javascript:showAbstract('Checkpoint and restart mechanisms have been widely used in large scientific simulation applications to make forward progress in case of failures. However, none of the prior works have considered the interaction of power-constraint with temperature, reliability, performance, and checkpointing interval. It is not clear how power-capping may affect optimal checkpointing interval. What are the involved reliability, performance, and energy trade-offs? In this paper, we develop a deep understanding about the interaction between power-capping and scientific applications using checkpoint\/restart as resilience mechanism, and propose a new model for the optimal checkpointing interval (OCI) under power-capping. Our study reveals several interesting, and previously unknown, insights about how power-capping affects the reliability, energy consumption, performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tang16power-capping.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#tang16power-capping\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Leonardo Bautista-Gomez, Ana Gainaru, Swann Perarnau, Devesh Tiwari, Saurabh Gupta, Franck Cappello, Christian Engelmann, and Marc Snir. <b>Reducing Waste in Extreme Scale Systems Through Introspective Analysis<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ipdps.org\" target=\"www.ipdps.org\">30th IEEE International Parallel and Distributed Processing Symposium (IPDPS) 2016<\/a><\/i>, pages 212-221, Chicago, IL, USA, May 23-27, 2016. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISSN 1530-2075. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/IPDPS.2016.100\" target=\"publication\">10.1109\/IPDPS.2016.100<\/a>. Acceptance rate 23.0% (114\/496). <a href=\"javascript:showAbstract('Resilience is an important challenge for extreme-scale  supercomputers. Today, failures in supercomputers are  assumed to be uniformly distributed in time. However, recent  studies show that failures in high-performance computing  systems are partially correlated in time, generating periods  of higher failure density. Our study of the failure logs of  multiple supercomputers show that periods of higher failure  density occur with up to three times more than the average.  We design a monitoring system that listens to hardware  events and forwards important events to the runtime to  detect those regime changes. We implement a runtime capable  of receiving notifications and adapt dynamically. In  addition, we build an analytical model to predict the gains  that such dynamic approach could achieve. We demonstrate that  in some systems, our approach can reduce the wasted time.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/bautista-gomez16reducing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/bautista-gomez16reducing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#bautista-gomez16reducing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Peer-reviewed Workshop Publications<\/h4>\n<ol>\n<li>Yawei Hui, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>A Comprehensive Informative Metric for Analyzing HPC System Status using the LogSCAN Platform<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc18.supercomputing.org\" target=\"sc18.supercomputing.org\">31st International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\">8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2018<\/a><\/i>, pages 29-38, Dallas, TX, USA, November 16, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-0222-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS.2018.00007\" target=\"publication\">10.1109\/FTXS.2018.00007<\/a>. Acceptance rate 45.0% (9\/20). <a href=\"javascript:showAbstract('Log processing by Spark and Cassandra-based ANalytics (LogSCAN) is a newly developed analytical platform that provides flexible and scalable data gathering, transformation and computation. One major challenge is to effectively summarize the status of a complex computer system, such as the Titan supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). Although there is plenty of operational and maintenance information collected and stored in real time, which may yield insights about short- and long-term system status, it is difficult to present this information in a comprehensive form. In this work, we present system information entropy (SIE), a newly developed metric that leverages the powers of traditional machine learning techniques and information theory. By compressing the multi-variant multi-dimensional event information recorded during the operation of the targeted system into a single time series of SIE, we demonstrate that the historical system status can be sensitively represented concisely and comprehensively. Given a sharp indicator as SIE, we argue that follow-up analytics based on SIE will reveal in-depth knowledge about system status using other sophisticated approaches, such as pattern recognition in the temporal domain or causality analysis incorporating extra independent metrics of the system.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18comprehensive2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hui18comprehensive2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hui18comprehensive2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf and Christian Engelmann. <b>Analyzing the Impact of System Reliability Events on Applications in the Titan Supercomputer<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc18.supercomputing.org\" target=\"sc18.supercomputing.org\">31st International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\">8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2018<\/a><\/i>, pages 39-48, Dallas, TX, USA, November 16, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-0222-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS.2018.00008\" target=\"publication\">10.1109\/FTXS.2018.00008<\/a>. Acceptance rate 45.0% (9\/20). <a href=\"javascript:showAbstract('Extreme-scale computing systems employ Reliability, Availability and Serviceability (RAS) mechanisms and infrastructure to log events from multiple system components. In this paper, we analyze RAS logs in conjunction with the  application placement and scheduling database, in order to  understand the impact of common RAS events on application performance. This study conducted on the records of about 2 million applications executed on Titan supercomputer provides important insights for system users, operators and computer science researchers. In this paper, we investigate the impact of RAS events on application performance and its variability by comparing cases where events are recorded with corresponding cases where no events are recorded. Such a statistical investigation is possible since we observed that system users tend to execute their applications multiple times. Our analysis reveals that most RAS events do impact application performance, although not always. We also find that different system components affect application performance differently. In particular, our investigation includes the following components: parallel file system, processor, memory, graphics processing units, system and user software issues. Our work establishes the importance of providing feedback to system users for increasing operational efficiency of extreme-scale systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18analyzing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18analyzing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18analyzing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Byung Hoon (Hoony) Park, Yawei Hui, Swen Boehm, Rizwan Ashraf, Christian Engelmann, and Christopher Layton. <b>A Big Data Analytics Framework for HPC Log Data: Three Case Studies Using the Titan Supercomputer Log<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster2018.github.io\" target=\"cluster2018.github.io\">19th IEEE International Conference on Cluster Computing (Cluster) 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/hpcmaspa2018\" target=\"sites.google.com\/site\/hpcmaspa2018\">5th Workshop on Monitoring and Analysis for High Performance Systems Plus Applications (HPCMASPA) 2018<\/a><\/i>, pages 571-579, Belfast, UK, September 10, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-8319-4. ISSN 2168-9253. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTER.2018.00073\" target=\"publication\">10.1109\/CLUSTER.2018.00073<\/a>. <a href=\"javascript:showAbstract('Reliability, availability and serviceability (RAS) logs of high performance computing (HPC) resources, when closely investigated in spatial and temporal dimensions, can provide invaluable information regarding system status, performance, and resource utilization. These data are often generated from multiple logging systems and sensors that cover many components of the system. The analysis of these data for finding persistent temporal and spatial insights faces two main difficulties: the volume of RAS logs makes manual inspection difficult and the unstructured nature and unique properties of log data produced by each subsystem adds another dimension of difficulty in identifying implicit correlation among recorded events. To address these issues, we recently developed a multi-user Big Data analytics framework for HPC log data at Oak Ridge National Laboratory (ORNL). This paper introduces three in-progress data analytics projects that leverage this framework to assess system status, mine event patterns, and study correlations between user applications and system events. We describe the motivation of each project and detail their workflows using three years of log data collected from ORNL&amp;#39;s Titan supercomputer.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/park18big.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/park18big.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#park18big\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Byung Hoon (Hoony) Park, Saurabh Hukerikar, Christian Engelmann, and Ryan Adamson. <b>Big Data Meets HPC Log Analytics: Scalable Approach to Understanding Systems at Extreme Scale<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster17.github.io\" target=\"cluster17.github.io\">18th IEEE International Conference on Cluster Computing (Cluster) 2017<\/a>: <a href=\"http:\/\/sites.google.com\/site\/hpcmaspa2017\" target=\"sites.google.com\/site\/hpcmaspa2017\">4th Workshop on Monitoring and Analysis for High Performance Systems Plus Applications (HPCMASPA) 2017<\/a><\/i>, pages 758-765, Honolulu, HI, USA, September 5, 2017. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-2327-5. ISSN 2168-9253. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTER.2017.113\" target=\"publication\">10.1109\/CLUSTER.2017.113<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s high-performance computing (HPC) systems are heavily instrumented generating logs containing information about abnormal events, such as critical conditions, faults, errors and failures, system resource utilization, and about the resource usage of user applications. These logs, once fully analyzed and correlated, can produce detailed information about the system health, root causes of failures, and analyze an application's interactions with the system, providing invaluable insights to domain scientists and system administrators. However, processing HPC logs requires deep understanding of hardware and software components at multiple layers of the system stack. Moreover, most log data is unstructured and voluminous, making it more difficult for scientists and engineers to analyze the data. With rapid increases in the scale and complexity of HPC systems, log data processing is becoming a big data challenge. This paper introduces a HPC log data analytics framework that is based on a distributed NoSQL database technology, which provides scalability and high availability, and the Apache Spark for rapid in-memory processing of log data. The framework enables the extraction of a range of information about the system so that system administrators and end users alike can obtain necessary insights for their specific needs. We describe our experience with using this framework to glean insights from the log data derived from the Titan supercomputer at the Oak Ridge National Laboratory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/park17big.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/park17big.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#park17big\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Peer-reviewed Conference Posters<\/h4>\n<ol>\n<li>Yawei Hui, Rizwan Ashraf, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>Real-Time Assessment of Supercomputer Status by a Comprehensive Informative Metric through Streaming Processing<\/b>. Poster at the  <a href=\"http:\/\/cci.drexel.edu\/bigdata\/bigdata2018\" target=\"cci.drexel.edu\/bigdata\/bigdata2018\">6th IEEE International Conference on Big Data (BigData) 2018<\/a>,  Seattle, WA, USA, December 10-13, 2018. <a href=\"javascript:showAbstract('Supercomputers are complex systems used to simulate, understand and solve real-world problems. In order to operate these systems efficiently and for the purpose of their maintainability, an accurate, concise, and timely determination of system status is crucial for its users and operators. However, this determination is challenging due to intricately connected heterogeneous software and hardware components, and due to sheer scale of such machines. In this poster, we demonstrate work-in-progress towards realization of a real-time monitoring framework for the 18,688-node Titan supercomputer at Oak Ridge Leadership Computing Facility (OLCF). Toward this end, we discuss the use of metrics which present a one-dimensional view of the system generating various types of information from 1000s of components and utilization statistics from 100s of user applications in near real-time. We demonstrate the efficacy of these metrics to understand and visualize raw log data generated by the system which otherwise may compose of 1000s of dimensions. We also demonstrate the architecture of proposed real-time stream processing framework which integrates, processes, analyzes, visualizes and stores system log data from an array of system components..');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18realtime.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hui18realtime\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>A Comprehensive Informative Metric for Summarizing HPC System Status<\/b>. Poster at the <a href=\"http:\/\/ldav.org\" target=\"ldav.org\">8th IEEE Symposium on Large Data Analysis and   Visualization<\/a> in conjunction with the   <a href=\"http:\/\/ieeevis.org\/year\/2018\" target=\"ieeevis.org\/year\/2018\">8th IEEE Vis 2018<\/a>,  Berlin, Germany, October 21, 2018. <a href=\"javascript:showAbstract('It remains a major challenge to effectively summarize and visualize in a comprehensive form the status of a complex computer system, such as the Titan supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). In the ongoing research highlighted in this poster, we present system information entropy (SIE), a newly developed system metric that leverages the powers of traditional machine learning techniques and information theory. By compressing the multi-variant multi-dimensional event information recorded during the operation of the targeted system into a single time series of SIE, we demonstrate that the historical system status can be sensitively summarized in form of SIE and visualized concisely and comprehensively.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18comprehensive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hui18comprehensive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>White Papers<\/h4>\n<ol>\n<li>Mingyan Li, Robert A. Bridges, Pablo Moriano, Christian Engelmann, Feiyi Wang, and Ryan Adamson. <b>Toward Effective Security\/Reliability Situational Awareness via Concurrent Security-or-Fault Analytics <\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/2021ascr-cybersecurity\" target=\"www.orau.gov\/2021ascr-cybersecurity\">ASCR Workshop on Cybersecurity and Privacy for Scientific  Computing Ecosystems<\/a><\/i>, November 3-5, 2021. <a href=\"javascript:showAbstract('Modern critical infrastructures (CI) and scientific computing ecosystems (SCE) are complex and vulnerable. The complexity of CI\/SCE, such as the distributed workload found across ASCR scientific computing facilities, does not allow for easy differentiation between emerging cyber security and reliability threats. It is also not easy to correctly identify the misbehaving systems. Sometimes, system failures are just caused by unintentional user misbehavior or actual hardware\/software reliability issues, but it may take some significant amount of time and effort to develop that understanding through root-cause analysis. On the security front, CI\/SCE are vital assets. They are prime targets of, and are vulnerable to, malicious cyber-attacks. Within DoE, inter-disciplinary and cross-facility collaboration (e.g., ORNL INTERSECT initiative, next-gen supercomputing OLCF6), traditional perimeter-based defense and demarcation line between malicious cyber-attacks and non-malicious system faults are blurring. Amidst realistic reliability and security threats, the ability to effectively distinguish between non-malicious faults and malicious attacks is critical not only in root cause identification but also in countermeasures generation. ');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/li21toward.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#li21toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Petar Radojkovic, Manolis Marazakis, Paul Carpenter, Reiley Jeyapaul, Dimitris Gizopoulos, Martin Schulz, Adria Armejach, Eduard Ayguade, Fran&ccedil;ois Bodin, Ramon Canal, Franck Cappello, Fabien Chaix, Guillaume Colin de Verdiere, Said Derradji, Stefano Di Carlo, Christian Engelmann, Ignacio Laguna, Miquel Moreto, Onur Mutlu, Lazaros Papadopoulos, Olly Perks, Manolis Ploumidis, Bezhad Salami, Yanos Sazeides, Dimitrios Soudris, Yiannis Sourdis, Per Stenstrom, Samuel Thibault, Will Toms, and Osman Unsal. <b>Towards Resilient EU HPC Systems: A Blueprint<\/b>. <i>White paper by the <a href=\"http:\/\/resilienthpc.eu\" target=\"resilienthpc.eu\">European HPC resilience initiative<\/a><\/i>, April 9, 2020. <a href=\"javascript:showAbstract('This document aims to spearhead a Europe-wide discussion on HPC system resilience and to help the European HPC community define best practices for resilience. We analyse a wide range of state-of-the-art resilience mechanisms and recommend the most effective approaches to employ in large-scale HPC systems. Our guidelines will be useful in the allocation of available resources, as well as guiding researchers and research funding towards the enhancement of resilience approaches with the highest priority and utility. Although our work is focussed on the needs of next generation HPC systems in Europe, the principles and evaluations are applicable globally. This document is the first output of the ongoing European HPC resilience initiative and it covers individual nodes in HPC systems, encompassing CPU, memory, intra-node interconnect and emerging FPGA-based hardware accelerators. With community support and feedback on this initial document, we will update the analysis and expand the scope to include other types of accelerators, as well as networks and storage.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/radojkovic20towards.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#radojkovic20towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Devesh Tiwari, Saurabh Gupta, and Christian Engelmann. <b>Lightweight, Actionable Analytical Tools Based on Statistical Learning for Efficient System Operations<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2016\" target=\"hpc.pnl.gov\/modsim\/2016\">Workshop on Modeling &#038; Simulation of Systems &#038; Applications (ModSim) 2016<\/a><\/i>, August 10-12, 2016. <a href=\"javascript:showAbstract('Modeling and simulation community has always relied on accurate and meaningful system data and parameters to drive analytical models and simulators. HPC systems continuously generate huge amount system event related data (e.g., system log, resource consumption log, RAS logs, power consumption logs), but meaningful interpretation and accuracy verification of such data is quite challenging. This talk offers a unique perspective and experience in demonstrating how modeling and simulation based research can actually be translated into production systems. We will discuss the short-term opportunities for modeling and simulation community to increase the impact and effectiveness of our analytical tools, &amp;#34;dos and don&amp;#39;ts&amp;#34;, long-term challenges and opportunities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tiwari16lightweight.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tiwari16lightweight.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tiwari16lightweight\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Marc Snir, and Robert W. Wisniewski, Jacob A. Abraham, Sarita V. Adve, Saurabh Bagchi, Pavan Balaji, Bill Carlson, Andrew A. Chien, Pedro Diniz, Christian Engelmann, Rinku Gupta, Fred Johnson, Jim Belak, Pradip Bose, Franck Cappello, Paul Coteus, Nathan A. Debardeleben, Mattan Erez, Saverio Fazzari, Al Geist, Sriram Krishnamoorthy, Sven Leyffer, Dean Liberty, Subhasish Mitra, Todd Munson, Rob Schreiber, Jon Stearley, and Eric Van Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i>Workshop report<\/i>, August 4-11, 2013. <a href=\"publications\/snir13addressing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir13addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Al Geist, Bob Lucas, Marc Snir, Shekhar Borkar, Eric Roman, Mootaz Elnozahy, Bert Still, Andrew Chien, Robert Clay, John Wu, Christian Engelmann, Nathan DeBardeleben, Rob Ross, Larry Kaplan, Martin Schulz, Mike Heroux, Sriram Krishnamoorthy, Lucy Nowell, Abhinav Vishnu, and Lee-Ann Talley. <b>U.S. Department of Energy Fault Management Workshop<\/b>. <i>Workshop report for the U.S. Department of Energy<\/i>, June 6, 2012. <a href=\"javascript:showAbstract('A Department of Energy (DOE) Fault Management Workshop was held on June 6, 2012 at the BWI Airport Marriot hotel in Maryland. The goals of this workshop were to: 1. Describe the required HPC resilience for critical DOE mission needs; 2. Detail what HPC resilience research is already being done at the DOE national laboratories and is expected to be done by industry or other groups; 3. Determine what fault management research is a priority for DOE&amp;#39;s Office of Science and National Nuclear Security Administration (NNSA) over the next five years; 4. Develop a roadmap for getting the necessary research accomplished in the timeframe when it will be needed by the large computing facilities across DOE.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/geist12department.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#geist12department\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Datasets<\/h4>\n<ol>\n<li>Woong Shin, Vladyslav Oles, Anna Schmedding, George Ostrouchov, Evgenia Smirni, Christian Engelmann, and Feiyi Wang. <b>OLCF Summit Supercomputer GPU Snapshots During Double-Bit Errors and Normal Operations<\/b>. <i><\/i>, April 20, 2023. DOI <a href=\"http:\/\/dx.doi.org\/10.13139\/OLCF\/1970187\" target=\"publication\">10.13139\/OLCF\/1970187<\/a>. <a href=\"javascript:showAbstract('As we move into the exascale era, the power and energy footprints of high-performance computing (HPC) systems have grown significantly larger. Due to the harsh power and thermal conditions the system, components are exposed to extreme operating conditions. Operation of such modern HPC systems requires deep insights into long term system behavior to maintain its efficiency as well as its longevity. To help the HPC community to gain such insights, we provide double-bit errors using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). The dataset relies on Nvidia XID records internally collected by GPU firmware at the time of failure occurrence, on the reboot-time logs of each Summit node, on node-level job scheduler records collected after each job termination, and on a 1Hz data rate from the baseboard management controllers (BMCs) of each Summit compute node using the OpenBMC event subscription protocol.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/shin23olcf.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#shin23olcf\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mallikarjun Shankar, George Ostrouchov, Don Maxwell, James Rogers, Rizwan Ashraf, and Christian Engelmann. <b>GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability<\/b>. <i><\/i>, September 2, 2020. DOI <a href=\"http:\/\/dx.doi.org\/10.13139\/ORNLNCCS\/1657202\" target=\"publication\">10.13139\/ORNLNCCS\/1657202<\/a>. <a href=\"publications\/shankar20gpu.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#shankar20gpu\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Talks and Lectures<\/h4>\n<ol>\n<li>Christian Engelmann. <b>Designing Smart and Resilient Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/conferences\/cm\/conference\/pp22\" target=\"www.siam.org\/conferences\/cm\/conference\/pp22\">20th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2022<\/a>, Seattle, WA, USA, February 23-26, 2022. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and in extrapolating this knowledge to future-generation systems. It also describes the path forward in machine-in-the-loop operational intelligence for smart computing systems, leveraging operational data analytics in a loop control that maximizes productivity and minimizes costs through adaptive autonomous operation for resilience.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22designing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann22designing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Faults, Errors and Failures in Extreme-Scale Supercomputers<\/b>. Keynote talk at the  <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\">14th  Workshop on Resiliency in High Performance Computing  (Resilience) in Clusters, Clouds, and Grids<\/a>, held in  conjunction with the <a href=\"http:\/\/europar2014.dcc.fc.up.pt\" target=\"europar2014.dcc.fc.up.pt\">27th European Conference on Parallel and Distributed  Computing (Euro-Par) 2021<\/a>, Lisbon, Portugal, August 30, 2021. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of reliability experiences with some of the largest supercomputers in the world and recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in these systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21faults.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21faults\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Resilience Problem in Extreme Scale Computing: Experiences and the Path Forward<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/conferences\/cm\/conference\/cse21\" target=\"www.siam.org\/conferences\/cm\/conference\/cse21\">SIAM Conference on Computational Science and Engineering (CSE) 2021<\/a>, Fort Worth, TX, USA, March 1-5, 2021. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of reliability experiences with some of the largest supercomputers in the world and recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in these systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Smart and Resilient Extreme-Scale Systems<\/b>. Invited talk at the  <a href=\"http:\/\/www.hipeac.net\/2021\/spring-virtual\/#\/program\/sessions\/7854\/\" target=\"www.hipeac.net\/2021\/spring-virtual\/#\/program\/sessions\/7854\/\">Workshop on Resilience in High Performance Computing  (RESILIENTHPC)<\/a>, held in conjunction with the  <a href=\"http:\/\/www.hipeac.net\/2021\" target=\"www.hipeac.net\/2021\">European Network on High-performance Embedded Architecture   and Compilation (HiPEAC) Conference 2021<\/a>, Budapest, Hungary, January 19, 2021. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and in extrapolating this knowledge to future-generation systems. It also describes the path forward in machine-in-the-loop operational intelligence for smart computing systems, leveraging operational data analytics in a loop control that maximizes productivity and minimizes costs through adaptive autonomous operation for resilience.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21smart.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21smart\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Resilience Problem in Extreme Scale Computing<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/conferences\/cm\/conference\/pp20\" target=\"www.siam.org\/conferences\/cm\/conference\/pp20\">19th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2020<\/a>, Seattle, WA, USA, February 12-15, 2020. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and in extrapolating this knowledge to future-generation systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann20resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann20resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience for Extreme Scale Systems: Understanding the Problem<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/cse19\/\" target=\"www.siam.org\/meetings\/cse19\/\">SIAM Conference on Computational Science and Engineering (CSE) 2019<\/a>, Spokane, WA, USA, February 25 &#8211; March 1, 2018. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of the Catalog project, which develops a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, this project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Characterizing Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/pasc18.pasc-conference.org\" target=\"pasc18.pasc-conference.org\">Platform for Advanced Scientific Computing (PASC) Conference 2018<\/a>, Basel, Switzerland, July 2-4, 2018. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18characterizing2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18characterizing2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Characterizing Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/iadac.github.io\/adac6\" target=\"iadac.github.io\/adac6\">6th Accelerated Data Analytics and Computing (ADAC) Institute Workshop<\/a>, Zurich, Switzerland, June 20-21, 2018. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18characterizing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18characterizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>A Catalog of Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/an17\/\" target=\"www.siam.org\/meetings\/an17\/\">SIAM Annual Meeting (AM) 2017<\/a>, Pittsburgh, PA, USA, July 10-14, 2017. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann17catalog2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann17catalog2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Characterizing Faults, Errors and Failures in Extreme-Scale Computing Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.isc-hpc.com\" target=\"www.isc-hpc.com\">International Supercomputing Conference (ISC) 2017<\/a>, Frankfurt am Main, Germany, June 16-22, 2017. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann17characterizing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann17characterizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>A Catalog of Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/icl.cs.utk.edu\/workshops\/scheduling2017\/\" target=\"icl.cs.utk.edu\/workshops\/scheduling2017\/\">12th Scheduling for Large Scale Systems Workshop (SLSSW) 2017<\/a>, Knoxville, TN, USA, May 24-26, 2017. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann17catalog.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann17catalog\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Missing High-Performance Computing Fault Model<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/pp16\/\" target=\"www.siam.org\/meetings\/pp16\/\">17th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2016<\/a>, Paris, France, April 12-15, 2016. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges. Resilience is one of the most important challenges. This talk will present recent work in developing the missing high-performance computing (HPC) fault model. This effort identifies, categorizes and models the fault, error and failure properties of today&amp;#39;s HPC systems. It develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current systems and extrapolates this knowledge to exascale HPC systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16missing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann16missing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience Challenges and Solutions for Extreme-Scale Supercomputing<\/b>. Invited talk at the <a href=\"http:\/\/www.usna.edu\" target=\"www.usna.edu\">United  States Naval Academy<\/a>, Annapolis, MD, USA, February 18, 2016. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2022) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2022). This talk provides an overview of recent and ongoing resilience research and development activities at Oak Ridge National Laboratory in advanced checkpoint storage architectures, process-level incremental checkpoint\/restart, proactive fault tolerance using prediction-triggered process or virtual machine migration, MPI process-level software redundancy, and soft-error injection tools to study the vulnerability of science applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16resilience2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann16resilience2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Toward A Fault Model And Resilience Design Patterns For Extreme Scale Systems<\/b>. Keynote talk at the  <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2015\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2015\">8th  Workshop on Resiliency in High Performance Computing  (Resilience) in Clusters, Clouds, and Grids<\/a>, held in  conjunction with the <a href=\"http:\/\/europar2014.dcc.fc.up.pt\" target=\"europar2014.dcc.fc.up.pt\">21st European Conference on Parallel and Distributed  Computing (Euro-Par) 2015<\/a>, Vienna, Austria, August 24-28, 2015. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2022) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2022). This talk provides an overview of two recently funded projects. The Characterizing Faults, Errors, and Failures in Extreme-Scale Systems project identifies, categorizes and models the fault, error and failure properties of US Department of Energy high-performance computing (HPC) systems. It develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current systems and extrapolate this knowledge to exascale HPC systems. The Resilience Design Patterns project will increase the ability of scientific applications to reach accurate solutions in a timely and efficient manner. Using a novel design pattern concept, it identifies and evaluates repeatedly occurring resilience problems and coordinates solutions throughout high-performance computing hardware and software.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann15toward.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann15toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>US Department of Energy (DOE) leadership computing facilities are in the process of deploying extreme-scale high-performance computing (HPC) systems with the long-range goal of building exascale systems that perform more than a quintillion (a billion billion) operations per second. More powerful computers mean researchers can simulate biological, chemical, and other physical interactions with an unprecedented&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":145,"menu_order":21,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-161","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/161","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=161"}],"version-history":[{"count":20,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/161\/revisions"}],"predecessor-version":[{"id":1288,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/161\/revisions\/1288"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/145"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=161"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}