{"id":16,"date":"2026-09-04T08:00:11","date_gmt":"2026-09-04T08:00:11","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=16"},"modified":"2026-09-05T01:06:18","modified_gmt":"2026-09-05T01:06:18","slug":"publications","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=16","title":{"rendered":"Publications"},"content":{"rendered":"<h4>\n<a name=\"journals\"><\/a>Peer-reviewed Journal Papers<\/h4>\n<ol>\n<li>Emmanuel Agullo, Mirco Altenbernd, Hartwig Anzt, Leonardo Bautista-Gomez, Tommaso Benacchio, Luca Bonaventura, Hans-Joachim Bungartz, Sanjay Chatterjee, Florina M. Ciorba, Nathan DeBardeleben, Daniel Drzisga, Sebastian Eibl, Christian Engelmann, Wilfried N. Gansterer, Luc Giraud, Dominik G&ouml;ddeke, Marco Heisig, Fabienne J&eacute;z&eacute;quel, Nils Kohl, Xiaoye Sherry Li, Romain Lion, Miriam Mehl, Paul Mycek, Michael Obersteiner, Enrique S. Quintana-Ort&iacute;, Francesco Rizzi, Ulrich R&uuml;de, Martin Schulz, Fred Fung, Robert Speck, Linda Stals, Keita Teranishi, Samuel Thibault, Dominik Th&ouml;nnes, Andreas Wagner, and Barbara Wohlmuth. <b>Resiliency in Numerical Algorithm Design for Extreme Scale Simulations<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 36, number 2, pages 251-285, March 1, 2022. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/10943420211055188\" target=\"publication\">10.1177\/10943420211055188<\/a>. <a href=\"javascript:showAbstract('This work is based on the seminar titled &amp;#39;Resiliency in Numerical Algorithm Design for Extreme Scale Simulations' held March 1-6, 2020 at Schloss Dagstuhl, that was attended by all the writers. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 hours on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 1023 floating- point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications, and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/agullo22resiliency.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#agullo22resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar, Saurabh Gupta, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, and Devesh Tiwari. <b>Study of Interconnect Errors, Network Congestion, and Applications Characteristics for Throttle Prediction on a Large Scale HPC System<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 153, pages 29-43, July 1, 2021. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2021.03.001\" target=\"publication\">10.1016\/j.jpdc.2021.03.001<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of peta ops. The performance of these systems depends on various components like processors, memory, and interconnect. Among  all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar21study.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar21study\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Amogh Katti, Giuseppe Di Fatta, Thomas Naughton, and Christian Engelmann. <b>Epidemic Failure Detection and Consensus for Extreme Parallelism<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 32, number 5, pages 729-743, September 1, 2018. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/1094342017690910\" target=\"publication\">10.1177\/1094342017690910<\/a>. <a href=\"javascript:showAbstract('Future extreme-scale high-performance computing systems will be required to work under frequent component failures. The MPI Forum&amp;#39;s User Level Failure Mitigation proposal has introduced an operation, MPI Comm shrink, to synchronize the alive processes on the list of failed processes, so that applications can continue to execute even in the presence of failures by adopting algorithm-based fault tolerance techniques. This MPI Comm shrink operation requires a failure detection and consensus algorithm. This paper presents three novel failure detection and consensus algorithms using Gossiping. The proposed algorithms were implemented and tested using the Extreme-scale Simulator. The results show that in all algorithms the number of Gossip cycles to achieve global consensus scales logarithmically with system size. The second algorithm also shows better scalability in terms of memory and network bandwidth usage and a perfect synchronization in achieving global consensus. The third approach is a three-phase distributed failure detection and consensus algorithm and provides consistency guarantees even in very large and extreme-scale systems while at the same time being memory and bandwidth efficient.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/katti18epidemic.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#katti18epidemic\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale<\/b>. <i><a href=\"http:\/\/superfri.org\/superfri\" target=\"superfri.org\/superfri\">Journal of Supercomputing Frontiers and Innovations (JSFI)<\/a><\/i>, volume 4, number 3, pages 4-42, October 1, 2017. <a href=\"http:\/\/www.susu.ru\/en\" target=\"www.susu.ru\/en\">South Ural State University Chelyabinsk, Russia<\/a>. ISSN 2409-6008. DOI <a href=\"http:\/\/dx.doi.org\/10.14529\/jsfi170301\" target=\"publication\">10.14529\/jsfi170301<\/a>. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires management of various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in HPC systems future systems are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. As a result, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods and metrics to investigate and evaluate resilience holistically in HPC systems that consider impact scope, handling coverage, and performance &amp; power efficiency across the system stack. Additionally, few of the current approaches are portable to newer architectures and software environments that will be deployed on future systems. In this paper, we develop a structured approach to the management of HPC resilience using the concept of resilience-based design patterns. A design pattern is a general repeatable solution to a commonly occurring problem. We identify the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. Each established solution is described in the form of a pattern that addresses concrete problems in the design of resilient systems. The complete catalog of resilience design patterns provides designers with reusable design elements. We also define a framework that enhances a designer&amp;#39;s understanding of the important constraints and opportunities for the design patterns to be implemented and deployed at various layers of the system stack. This design framework may be used to establish mechanisms and interfaces to coordinate flexible fault management across hardware and software components. The framework also supports optimization of the cost-benefit trade-offs among performance, resilience, and power consumption. The overall goal of this work is to enable a systematic methodology for the design and evaluation of resilience technologies in extreme-scale HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17resilience.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar17resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A New Deadlock Resolution Protocol and Message Matching Algorithm for the Extreme-scale Simulator<\/b>. <i><a href=\"http:\/\/onlinelibrary.wiley.com\/journal\/10.1002\/(ISSN)1532-0634\" target=\"onlinelibrary.wiley.com\/journal\/10.1002\/(ISSN)1532-0634\">Concurrency and Computation: Practice and Experience<\/a><\/i>, volume 28, number 12, pages 3369-3389, August 1, 2016. <a href=\"http:\/\/www.wiley.com\" target=\"www.wiley.com\">John Wiley &#038; Sons, Inc.<\/a>. ISSN 1532-0634. DOI <a href=\"http:\/\/dx.doi.org\/10.1002\/cpe.3805\" target=\"publication\">10.1002\/cpe.3805<\/a>. <a href=\"javascript:showAbstract('Investigating the performance of parallel applications at scale on future high-performance computing (HPC) architectures and the performance impact of different HPC architecture choices is an important component of HPC hardware\/software co-design. The Extreme-scale Simulator (xSim) is a simulation toolkit for investigating the performance of parallel applications at scale. xSim scales to millions of simulated Message Passing Interface (MPI) processes. The xSim toolkit strives to limit simulation overheads in order to maintain performance and productivity criteria. This paper documents two improvements to xSim: (1) a new deadlock resolution protocol to reduce the parallel discrete event simulation overhead, and (2) a new simulated MPI message matching algorithm to reduce the oversubscription management cost. These enhancements resulted in significant performance improvements. The simulation overhead for running the NAS Parallel Benchmark suite dropped from 1,020% to 238% for the conjugate gradient (CG) benchmark and 102% to 0% for the embarrassingly parallel (EP) benchmark. Additionally, the improvements were beneficial for reducing overheads in the highly accurate simulation mode of xSim, which is useful for resilience investigation studies for tracking intentional MPI process failures. In the highly accurate mode, the simulation overhead was reduced from 37,511% to 13,808% for CG and from 3,332% to 204% for EP.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16new.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann16new\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Marc Snir, Robert W. Wisniewski, Jacob A. Abraham, Sarita V. Adve, Saurabh Bagchi, Pavan Balaji, Jim Belak, Pradip Bose, Franck Cappello, Bill Carlson, Andrew A. Chien, Paul Coteus, Nathan A. Debardeleben, Pedro Diniz, Christian Engelmann, Mattan Erez, Saverio Fazzari, Al Geist, Rinku Gupta, Fred Johnson, Sriram Krishnamoorthy, Sven Leyffer, Dean Liberty, Subhasish Mitra, Todd Munson, Rob Schreiber, Jon Stearley, and Eric Van Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 28, number 2, pages 127-171, May 1, 2014. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/1094342014522573\" target=\"publication\">10.1177\/1094342014522573<\/a>. <a href=\"javascript:showAbstract('We present here a report produced by a workshop on  Addressing failures in exascale computing&amp;#39; held in Park City,  Utah, 4-11 August 2012. The charter of this workshop was to  establish a common taxonomy about resilience across all the  levels in a computing system, discuss existing knowledge on  resilience across the various hardware and software layers  of an exascale system, and build on those results, examining  potential solutions from both a hardware and software  perspective and focusing on a combined approach. The workshop brought together participants with expertise in  applications, system software, and hardware; they came from  industry, government, and academia, and their interests ranged  from theory to implementation. The combination allowed broad  and comprehensive discussions and led to this document, which  summarizes and builds on those discussions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/snir14addressing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir14addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Scaling To A Million Cores And Beyond: Using Light-Weight Simulation to Understand The Challenges Ahead On The Road To Exascale<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/fgcs\" target=\"www.elsevier.com\/locate\/fgcs\">Future Generation Computer Systems (FGCS)<\/a><\/i>, volume 30, number 0, pages 59-65, January 1, 2014. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0167-739X. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.future.2013.04.014\" target=\"publication\">10.1016\/j.future.2013.04.014<\/a>. <a href=\"javascript:showAbstract('As supercomputers scale to 1,000 PFlop\/s over the next decade, investigating the performance of parallel applications at scale on future architectures and the performance impact of different architecture choices for high-performance computing (HPC) hardware\/software co-design is crucial. This paper summarizes recent efforts in designing and implementing a novel HPC hardware\/software co-design toolkit. The presented Extreme-scale Simulator (xSim) permits running an HPC application in a controlled environment with millions of concurrent execution threads while observing its performance in a simulated extreme-scale HPC system using architectural models and virtual timing. This paper demonstrates the capabilities and usefulness of the xSim performance investigation toolkit, such as its scalability to 2^27 simulated Message Passing Interface (MPI) ranks on 960 real processor cores, the capability to evaluate the performance of different MPI collective communication algorithms, and the ability to evaluate the performance of a basic Monte Carlo application with different architectural parameters.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13scaling.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann13scaling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Proactive Process-Level Live Migration and Back Migration in HPC Environments<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 72, number 2, pages 254-267, February 1, 2012. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2011.10.009\" target=\"publication\">10.1016\/j.jpdc.2011.10.009<\/a>. <a href=\"javascript:showAbstract('As the number of nodes in high-performance computing environments keeps increasing, faults are becoming common place. Reactive fault tolerance (FT) often does not scale due to massive I\/O requirements and relies on manual job resubmission. This work complements reactive with proactive FT at the process level. Through health monitoring, a subset of node failures can be anticipated when one&amp;#39;s health deteriorates. A novel process-level live migration mechanism supports continued execution of applications during much of process migration. This scheme is integrated into an MPI execution environment to transparently sustain health-inflicted node failures, which eradicates the need to restart and requeue MPI jobs. Experiments indicate that 1-6.5 s of prior warning are required to successfully trigger live process migration while similar operating system virtualization mechanisms require 13-24 s. This self-healing approach complements reactive FT by nearly cutting the number of checkpoints in half when 70% of the faults are handled proactively. The work also provides a novel back migration approach to eliminate load imbalance or bottlenecks caused by migrated tasks. Experiments indicate the larger the amount of outstanding execution, the higher the benefit due to back migration.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang12proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#wang12proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, Christian Engelmann, and Hong H. Ong. <b>System-Level Virtualization Research at Oak Ridge National Laboratory<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/fgcs\" target=\"www.elsevier.com\/locate\/fgcs\">Future Generation Computer Systems (FGCS)<\/a><\/i>, volume 26, number 3, pages 304-307, March 1, 2010. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0167-739X. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.future.2009.07.001\" target=\"publication\">10.1016\/j.future.2009.07.001<\/a>. <a href=\"javascript:showAbstract('System-level virtualization is today enjoying a rebirth as a technique to effectively share what were then considered large computing resources to subsequently fade from the spotlight as individual workstations gained in popularity with a one machine - one user approach. One reason for this resurgence is that the simple workstation has grown in capability to rival that of anything available in the past. Thus, computing centers are again looking at the price\/performance benefit of sharing that single computing box via server consolidation. However, industry is only concentrating on the benefits of using virtualization for server consolidation (enterprise computing) whereas our interest is in leveraging virtualization to advance high-performance computing (HPC). While these two interests may appear to be orthogonal, one consolidating multiple applications and users on a single machine while the other requires all the power from many machines to be dedicated solely to its purpose, we propose that virtualization does provide attractive capabilities that may be exploited to the benefit of HPC interests. This does raise the two fundamental questions of: is the concept of virtualization (a machine sharing technology) really suitable for HPC and if so, how does one go about leveraging these virtualization capabilities for the benefit of HPC. To address these questions, this document presents ongoing studies on the usage of system-level virtualization in a HPC context. These studies include an analysis of the benefits of system-level virtualization for HPC, a presentation of research efforts based on virtualization for system availability, and a presentation of research efforts for the management of virtual systems. The basis for this document was material presented by Stephen L. Scott at the Collaborative and Grid Computing Technologies meeting held in Cancun, Mexico on April 12-14, 2007.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott10system.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott10system\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xubin (Ben) He, Li Ou, Christian Engelmann, Xin Chen, and Stephen L. Scott. <b>Symmetric Active\/Active Metadata Service for High Availability Parallel File Systems<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 69, number 12, pages 961-973, December 1, 2009. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2009.08.004\" target=\"publication\">10.1016\/j.jpdc.2009.08.004<\/a>. <a href=\"javascript:showAbstract('High availability data storage systems are critical for many applications as research and business become more data-driven. Since metadata management is essential to system availability, multiple metadata services are used to improve the availability of distributed storage systems. Past research focused on the active\/standby model, where each active service has at least one redundant idle backup. However, interruption of service and even some loss of service state may occur during a fail-over depending on the used replication technique. In addition, the replication overhead for multiple metadata services can be very high. The research in this paper targets the symmetric active\/active replication model, which uses multiple redundant service nodes running in virtual synchrony. In this model, service node failures do not cause a fail-over to a backup and there is no disruption of service or loss of service state. We further discuss a fast delivery protocol to reduce the latency of the needed total order broadcast. Our prototype implementation shows that metadata service high availability can be achieved with an acceptable performance trade-off using our symmetric active\/active metadata service solution.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/he09symmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#he09symmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xubin (Ben) He, Li Ou, Martha J. Kosa, Stephen L. Scott, and Christian Engelmann. <b>A Unified Multiple-Level Cache for High Performance Cluster Storage Systems<\/b>. <i><a href=\"http:\/\/www.inderscience.com\/browse\/index.php?journalcode=ijhpcn\" target=\"www.inderscience.com\/browse\/index.php?journalcode=ijhpcn\">International Journal of High Performance Computing and Networking (IJHPCN)<\/a><\/i>, volume 5, number 1-2, pages 97-109, November 14, 2007. <a href=\"http:\/\/www.inderscience.com\" target=\"www.inderscience.com\">Inderscience Publishers, Geneve, Switzerland<\/a>. ISSN 1740-0562. DOI <a href=\"http:\/\/dx.doi.org\/10.1504\/IJHPCN.2007.015768\" target=\"publication\">10.1504\/IJHPCN.2007.015768<\/a>. <a href=\"javascript:showAbstract('Highly available data storage for high-performance computing is becoming increasingly more critical as high-end computing systems scale up in size and storage systems are developed around network-centered architectures. A promising solution is to harness the collective storage potential of individual workstations much as we harness idle CPU cycles due to the excellent price\/performance ratio and low storage usage of most commodity workstations. For such a storage system, metadata consistency is a key issue assuring storage system availability as well as data reliability. In this paper, we present a decentralized metadata management scheme that improves storage availability without sacrificing performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/he07unified.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#he07unified\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Symmetric Active\/Active High Availability for High-Performance Computing System Services<\/b>. <i><a href=\"http:\/\/www.jcomputers.us\" target=\"www.jcomputers.us\">Journal of Computers (JCP)<\/a><\/i>, volume 1, number 8, pages 43-54, December 1, 2006. <a href=\"http:\/\/www.jcomputers.us\" target=\"www.jcomputers.us\">Academy Publisher, Oulu, Finland<\/a>. ISSN 1796-203X. DOI <a href=\"http:\/\/dx.doi.org\/10.4304\/jcp.1.8.43-54\" target=\"publication\">10.4304\/jcp.1.8.43-54<\/a>. <a href=\"javascript:showAbstract('This work aims to pave the way for high availability in high-performance computing (HPC) by focusing on efficient redundancy strategies for head and service nodes. These nodes represent single points of failure and control for an entire HPC system as they render it inaccessible and unmanageable in case of a failure until repair. The presented approach introduces two distinct replication methods, internal and external, for providing symmetric active\/active high availability for multiple redundant head and service nodes running in virtual synchrony utilizing an existing process group communication system for service group membership management and reliable, totally ordered message delivery. Resented results of a prototype implementation that offers symmetric active\/active replication for HPC job and resource management using external replication show that the highest level of availability can be provided with an acceptable performance trade-off.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06symmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann06symmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, David E. Bernholdt, Narasimha R. Gottumukkala, Chokchai (Box) Leangsuksun, Jyothish Varma, Chao Wang, Frank Mueller, Aniruddha G. Shet, and Ponnuswamy (Saday) Sadayappan. <b>MOLAR: Adaptive Runtime Support for High-End Computing Operating and Runtime Systems<\/b>. <i><a href=\"http:\/\/www.sigops.org\/osr.html\" target=\"www.sigops.org\/osr.html\">ACM SIGOPS Operating Systems Review (OSR)<\/a><\/i>, volume 40, number 2, pages 63-72, April 1, 2006. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISSN 0163-5980. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1131322.1131337\" target=\"publication\">10.1145\/1131322.1131337<\/a>. <a href=\"javascript:showAbstract('MOLAR is a multi-institutional research effort that concentrates on adaptive, reliable, and efficient operating and runtime system (OS\/R) solutions for ultra-scale, high-end scientific computing on the next generation of supercomputers. This research addresses the challenges outlined in FAST-OS (forum to address scalable technology for runtime and operating systems) and HECRTF (high-end computing revitalization task force) activities by exploring the use of advanced monitoring and adaptation to improve application performance and predictability of system interruptions, and by advancing computer reliability, availability and serviceability (RAS) management systems to work cooperatively with the OS\/R to identify and preemptively resolve system issues. This paper describes recent research of the MOLAR team in advancing RAS for high-end computing OS\/Rs.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06molar.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann06molar\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"conferences\"><\/a>Peer-reviewed Conference Papers<\/h4>\n<ol>\n<li>Pedro Valero-Lara, Aaron Young, Thomas Naughton, Christian Engelmann, Al Geist, Jeffrey S. Vetter, Keita Teranishi, and William F. Godoy. <b>ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.sca-hpcasia2026.jp\" target=\"www.sca-hpcasia2026.jp\">Supercomputing Asia \/ International Conference on High Performance Computing in the Asia-Pacific Region (SCA\/HPCAsia) 2026<\/a><\/i>, pages 19-30, Osaka, Japan, January 26-29, 2026. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 979-8-4007-2067-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3773656.3773659\" target=\"publication\">10.1145\/3773656.3773659<\/a>. Acceptance rate 36.6% (37\/101). <a href=\"javascript:showAbstract('The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually--especially applying a proper domain decomposition and communication pattern--is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)-based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4x boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/valero-lara26chatmpi.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#valero-lara26chatmpi\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Kibaek Kim, Krishnan Raghavan, Olivera Kotevska, Matthieu Dorier, Ravi Madduri, Minseok Ryu, Todd Munson, Rob Ross, Thomas Flynn, Ai Kagawa, Byung-Jun Yoon, Christian Engelmann, and Farzad Yousefian. <b>Privacy-Preserving Federated Learning for Science: Challenges and Research Directions<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ieeebigdata2024.github.io\" target=\"ieeebigdata2024.github.io\">12th IEEE International Conference on Big Data  (BigData) 2024<\/a><\/i>, pages 7849-7853, Washington, DC, USA, December 15-18, 2024. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 979-8-3503-6249-7. ISSN 2639-1589. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/BigData62323.2024.10825853\" target=\"publication\">10.1109\/BigData62323.2024.10825853<\/a>. Acceptance rate 18.5% (122\/661). <a href=\"javascript:showAbstract('This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kim24privacy.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kim24privacy\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Vladyslav Oles, Anna Schmedding, George Ostrouchov, Woong Shi, Evgenia Smirni, and Christian Engelmann. <b>Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ics2024.github.io\" target=\"ics2024.github.io\">38th ACM International Conference on Supercomputing  (ICS) 2024<\/a><\/i>, pages 188-200, Kyoto, Japan, June 4-7, 2024. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 979-8-4007-0610-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3650200.3656615\" target=\"publication\">10.1145\/3650200.3656615<\/a>. Acceptance rate 36.0% (45\/125). <a href=\"publications\/oles24understanding.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/oles24understanding.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#oles24understanding\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Suhas Somnath. <b>Science Use Case Design Patterns for Autonomous Experiments<\/b>. In <i>Proceedings of the <a href=\"http:\/\/europlop.net\" target=\"europlop.net\">28th European Conference on Pattern Languages of Programs (EuroPLoP) 2023<\/a><\/i>, pages 1-14, Kloster Irsee, Germany, July 5-9, 2023. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 979-8-4007-0040-8. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3628034.3628060\" target=\"publication\">10.1145\/3628034.3628060<\/a>. <a href=\"javascript:showAbstract('Connecting scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing (HPC) center enables autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) Open Architecture enables science break- throughs using intelligent networked systems, instruments and facilities with a federated hardware\/software architecture for the laboratory of the future. It relies on a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture. This paper introduces the science use case design patterns of the INTERSECT Architecture. It describes the overall background, the involved terminology and concepts, and the pattern format and classification. It further offers an overview of the 12 defined patterns and 4 examples of patterns of 2 different pattern classes. It also provides insight into building solutions from these patterns. The target audience are computer, computational, instrument and domain science experts working in the field of autonomous experiments.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23science.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann23science\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Olga Kuchar, Swen Boehm, Michael J. Brim, Thomas Naughton, Suhas Somnath, Scott Atchley, Jack Lange, Ben Mintz, and Elke Arenholz. <b>The INTERSECT Open Federated Architecture for the Laboratory of the Future<\/b>. In <i>Communications in Computer and Information Science (CCIS): Accelerating Science and Engineering Discoveries Through Integrated Research Infrastructure for Experiment, Big Data, Modeling and Simulation. <a href=\"http:\/\/smc.ornl.gov\" target=\"smc.ornl.gov\">18th Smoky Mountains Computational Sciences &#038; Engineering Conference (SMC) 2022<\/a><\/i>, pages 173-190, August 24-25, 2022. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer, Cham<\/a>. ISBN 978-3-031-23605-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-031-23606-8_11\" target=\"publication\">10.1007\/978-3-031-23606-8_11<\/a>. Acceptance rate 32.4% (24\/74). <a href=\"javascript:showAbstract('A federated instrument-to-edge-to-center architecture is needed to autonomously collect, transfer, store, process, curate, and archive scientific data and reduce human-in-the-loop needs with (a) common interfaces to leverage community and custom software, (b) pluggability to permit adaptable solutions, reuse, and digital twins, and (c) an open standard to enable adoption by science facilities world-wide. The INTERSECT Open Architecture enables science breakthroughs using intelligent networked systems, instruments and facilities with autonomous experiments, &amp;#34;self-driving&amp;#34; laboratories, smart manufacturing and \\glsAI driven design, discovery and evaluation. It creates an open federated architecture for the laboratory of the future using a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22intersect.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann22intersect.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann22intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>PLEXUS: A Pattern-Oriented Runtime System Architecture for Resilient Extreme-Scale High-Performance Computing Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/prdc.dependability.org\/PRDC2020\" target=\"prdc.dependability.org\/PRDC2020\">25th IEEE Pacific Rim International Symposium on  Dependable Computing (PRDC) 2020<\/a><\/i>, pages 31-39, Perth, Australia, December 1-4, 2020. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-8004-5. ISSN 1555-094X. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PRDC50213.2020.00014\" target=\"publication\">10.1109\/PRDC50213.2020.00014<\/a>. Acceptance rate 40.9% (18\/44). <a href=\"javascript:showAbstract('For high-performance computing (HPC) system designers and users, meeting the myriad challenges of next-generation exascale supercomputing systems requires rethinking their approach to application and system software design. Among these challenges, providing resiliency and stability to the scientific applications in the presence of high fault rates requires new approaches to software architecture and design. As HPC systems become increasingly complex, they require intricate solutions for detection and mitigation for various modes of faults and errors that occur in these large-scale systems, as well as solutions for failure recovery. These resiliency solutions often interact with and affect other system properties, including application scalability, power and energy efficiency. Therefore, resilience solutions for HPC systems must be thoughtfully engineered and deployed. In previous work, we developed the concept of resilience design patterns, which consist of templated solutions based on well-established techniques for detection, mitigation and recovery. In this paper, we use these patterns as the foundation to propose new approaches to designing runtime systems for HPC systems. The instantiation of these patterns within a runtime system enables flexible and adaptable end-to-end resiliency solutions for HPC environments. The paper describes the architecture of the runtime system, named Plexus, and the strategies for dynamically composing and adapting pattern instances under runtime control. This runtime-based approach enables actively balancing the cost-benefit trade-off between performance overhead and protection coverage of the resilience solutions. Based on a prototype implementation of PLEXUS, we demonstrate the resiliency and performance gains achieved by the pattern-based runtime system for a parallel linear solver application.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar20plexus.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar20plexus\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>George Ostrouchov, Don Maxwell, Rizwan Ashraf, Christian Engelmann, Mallikarjun Shankar, and James Rogers. <b>GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc20.supercomputing.org\" target=\"sc20.supercomputing.org\">33rd IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2020<\/a><\/i>, pages 41:1-14, Atlanta, GA, USA, November 15-20, 2020. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 9781728199986. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SC41405.2020.00045\" target=\"publication\">10.1109\/SC41405.2020.00045<\/a>. Acceptance rate 25.1% (95\/378). <a href=\"javascript:showAbstract('The Cray XK7 Titan was the top supercomputer system in the world for a very long time and remained critically important throughout its nearly seven year life. It was also a very interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three very significant rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 operation years in the GPU lifetimes, which correspond to Titan&amp;#39;s 6 year long productive period after an initial break-in period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the system cooling architecture and job scheduling. In addition to describing some of the system history, the data collection, data cleaning, and our analysis of the data, we provide reliability recommendations for designing future state of the art supercomputing systems and their operation. We make the data and our analysis codes publicly available.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ostrouchov20gpu.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ostrouchov20gpu.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ostrouchov20gpu\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Haewon Jeong, Yaoqing Yang, Christian Engelmann, Vipul Gupta, Tze Meng Low, Pulkit Grover, Viveck Cadambe, and Kannan Ramchandran. <b>3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.euro-par.org\" target=\"www.euro-par.org\">26th European Conference on Parallel and Distributed Computing (Euro-Par) 2020<\/a><\/i>, pages 392-407, Warsaw, Poland, August 24-28, 2020. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-030-57674-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-57675-2_25\" target=\"publication\">10.1007\/978-3-030-57675-2_25<\/a>. Acceptance rate 24.5% (39\/159). <a href=\"javascript:showAbstract('In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that is communication efficient and achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for distributed matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires 50% less redundancy than replication, while the overhead in execution time is only about 5-10%.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/jeong203d.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/jeong203d.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#jeong203d\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar, Saurabh Gupta, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, and Devesh Tiwari. <b>Understanding and Analyzing Interconnect Errors and Network Congestion on a Large Scale HPC System<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.dsn.org\" target=\"www.dsn.org\">48th IEEE\/IFIP International Conference on Dependable  Systems and Networks (DSN) 2018<\/a><\/i>, pages 107-114, Luxembourg City, Luxembourg, June 25-28, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-5596-2. ISSN 2158-3927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DSN.2018.00023\" target=\"publication\">10.1109\/DSN.2018.00023<\/a>. Acceptance rate 27.2% (62\/228). <a href=\"javascript:showAbstract('Today&amp;#39;s High Performance Computing (HPC) systems are capable of delivering performance in the order of petaflops due to the fast computing devices, network interconnect, and back-end storage systems. In particular, interconnect resilience and congestion resolution methods have a major impact on the overall interconnect and application performance. This is especially true for scientific applications running multiple processes on different compute nodes as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks state-of-practice experience reports that detail how different interconnect errors and congestion events occur on large-scale HPC systems. Therefore, in this paper, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors and congestion events. We also study the interaction between interconnect, errors, network congestion and application characteristics.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar18understanding.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar18understanding\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Bin Nie, Ji Xue, Saurabh Gupta, Tirthak Patel, Christian Engelmann, Evgenia Smirni, and Devesh Tiwari. <b>Machine Learning Models for GPU Error Prediction in a Large Scale HPC System<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.dsn.org\" target=\"www.dsn.org\">48th IEEE\/IFIP International Conference on Dependable  Systems and Networks (DSN) 2018<\/a><\/i>, pages 95-106, Luxembourg City, Luxembourg, June 25-28, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-5596-2. ISSN 2158-3927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DSN.2018.00022\" target=\"publication\">10.1109\/DSN.2018.00022<\/a>. Acceptance rate 27.2% (62\/228). <a href=\"javascript:showAbstract('Recently, GPUs have been widely deployed on large-scale HPC systems to provide powerful computational capability for scientific applications from various domains. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative. Therefore, in this paper, we firstly study the conditions that trigger GPU errors with six-month trace data collected from a large-scale operational HPC system. Then, we resort to machine learning techniques to predict the occurrence of GPU errors, by taking advantage of the temporal and spatial dependency of the collected data. As discussed in the evaluation section, the prediction framework is robust and accurate under different workloads.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/nie18machine.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#nie18machine\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf, Saurabh Hukerikar, and Christian Engelmann. <b>Pattern-based Modeling of Multiresilience Solutions for High-Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icpe2018.spec.org\" target=\"icpe2018.spec.org\">9th ACM\/SPEC International Conference on Performance Engineering (ICPE) 2018<\/a><\/i>, pages 80-87, Berlin, Germany, April 9-13, 2018. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-5095-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3184407.3184421\" target=\"publication\">10.1145\/3184407.3184421<\/a>. Acceptance rate 23.7% (14\/59). <a href=\"javascript:showAbstract('Resiliency is the ability of large-scale high-performance computing (HPC) applications to gracefully handle different types of errors, and recover from failures. In this paper, we propose a pattern-based approach to constructing multiresilience solutions. Using resilience patterns, we evaluate the performance and reliability characteristics of detection, containment and mitigation techniques for transient errors that cause silent data corruptions and techniques for fail-stop errors that result in process failures. We demonstrate the design and implementation of the resilience techniques across multiple layers of the system stack such that they are integrated to work together to achieve resiliency to different error types in a highly performance-effcient manner.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18pattern-based.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18pattern-based.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18pattern-based\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf, Saurabh Hukerikar, and Christian Engelmann. <b>Shrink or Substitute: Handling Process Failures in HPC Systems using In-situ Recovery<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2018.org\" target=\"www.pdp2018.org\">26th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2018<\/a><\/i>, pages 178-185, Cambridge, UK, March 21-23, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-4975-6. ISSN 2377-5750. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP2018.2018.00032\" target=\"publication\">10.1109\/PDP2018.2018.00032<\/a>. Acceptance rate 29.3% (27\/92). <a href=\"javascript:showAbstract('Efficient utilization of today&amp;#39;s high-performance computing (HPC) systems with many, complex software and hardware components requires that the HPC applications are designed to tolerate process failures at runtime. With low mean-time-to-failure (MTTF) of current and future HPC systems, long running simulations on these systems requires capabilities for gracefully handling process failures by the applications themselves. In this paper, we explore the use of fault tolerance extensions to Message Passing Interface (MPI) called user-level failure mitigation (ULFM) for handling process failures without the need to discard the progress made by the application. We explore two alternative recovery strategies, which use ULFM along with application-driven in-memory checkpointing. In the first case, the application is recovered with only the surviving processes, and in the second case, spares are used to replace the failed processes, such that the original configuration of the application is restored. Our experimental results demonstrate that graceful degradation is a viable alternative for recovery in environments where spares may not be available.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18shrink.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18shrink.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18shrink\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Gupta, Tirthak Patel, Christian Engelmann, and Devesh Tiwari. <b>Failures in Large Scale Systems: Long-term Measurement, Analysis, and Implications<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc17.supercomputing.org\" target=\"sc17.supercomputing.org\">30th IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2017<\/a><\/i>, pages 44:1-44:12, Denver, CO, USA, November 12-17, 2017. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-5114-0. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3126908.3126937\" target=\"publication\">10.1145\/3126908.3126937<\/a>. Acceptance rate 18.7% (61\/327). <a href=\"javascript:showAbstract('Resilience is one of the key challenges in maintaining high efficiency of future extreme scale supercomputers. Unfortunately, field-data based reliability studies are far in between and not exhaustive. Most HPC researchers and system practitioners still rely on outdated studies to understand HPC reliability characteristics and plan for future HPC systems. While the complexity of managing system reliability has increased, the public knowledge sharing about lessons learned from HPC centers has not increased in the same proportion. To bridge this gap, in this work, we compare and contrast the reliability characteristics of multiple large-scale HPC production systems, and discuss new take-aways and con rm previous findings which continue to be valid.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/gupta17failures.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/gupta17failures.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#gupta17failures\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Bin Nie, Ji Xue, Saurabh Gupta, Christian Engelmann, Evgenia Smirni, and Devesh Tiwari. <b>Characterizing Temperature, Power, and Soft-Error Behaviors in Data Center Systems: Insights, Challenges, and Opportunities<\/b>. In <i>Proceedings of the <a href=\"http:\/\/mascots2017.cs.ucalgary.ca\" target=\"mascots2017.cs.ucalgary.ca\">25th IEEE International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS) 2017<\/a><\/i>, pages 22-31, Banff, AB, Canada, September 20-22, 2017. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-2764-8. ISSN 2375-0227. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/MASCOTS.2017.12\" target=\"publication\">10.1109\/MASCOTS.2017.12<\/a>. Acceptance rate 30.95% (26\/84). <a href=\"javascript:showAbstract('GPUs have become part of the mainstream high performance computing facilities that increasingly require more computational power to simulate physical phenomena quickly and accurately. However, GPU nodes also consume significantly more power than traditional CPU nodes, and high power consumption introduces new system operation challenges, including increased temperature, power\/cooling cost, and lower system reliability. This paper explores how power consumption and temperature characteristics affect reliability, provides insights into what are the implications of such understanding, and how to exploit these insights toward predicting GPU errors using neural networks.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/nie17characterizing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#nie17characterizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>A Pattern Language for High-Performance Computing Resilience<\/b>. In <i>Proceedings of the <a href=\"http:\/\/europlop.net\" target=\"europlop.net\">22nd European Conference on Pattern Languages of Programs (EuroPLoP) 2017<\/a><\/i>, pages 12:1-12:16, Kloster Irsee, Germany, July 12-16, 2017. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-4848-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3147704.3147718\" target=\"publication\">10.1145\/3147704.3147718<\/a>. <a href=\"javascript:showAbstract('High-performance computing systems (HPC) provide powerful capabilities for modeling and simulation, and data analytics in a broad class of computational problems in a variety of scientific and engineering domains. HPC designs are undergoing rapid changes in the hardware architectures and the software environment as the community pursues increasingly capable HPC systems. Among the key challenges for future generations of HPC systems is the ensuring efficient and correct operation despite the occurrence of faults or defects in system components that can cause errors and failures in a HPC system. Such events affect the correctness of the scientific applications, or may lead to their untimely termination. Future generations of HPC systems will consist of millions of compute, memory and storage components and the growing complexity of these computing behemoths increases the chances that a single fault event will cascade across the machine and bring down the entire system. Design patterns capture the essential techniques that are employed to solve recurring problems in the design of resilient computing systems. However, the complexity of modern HPC systems as well as the various challenges of future generations of systems requires consideration to numerous aspects and optimization principles, such as the impact of a resilience solution on the performance and power consumption. We present a pattern language for engineering resilience solutions. The language is targeted at hardware and software designers as well as the users and operators of HPC systems. The patterns are intended to develop complete resilience solutions that have different efficiency and complexity characteristics, which may be deployed at design time or runtime to ensure that HPC systems are able to deal with various types of faults, errors and failures.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17pattern.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar17pattern\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mahesh Lagadapati, Frank Mueller, and Christian Engelmann. <b>Benchmark Generation and Simulation at Extreme Scale<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ds-rt.com\/2016\" target=\"ds-rt.com\/2016\">20th IEEE\/ACM International Symposium on Distributed Simulation and Real Time Applications (DS-RT) 2016<\/a><\/i>, pages 9-18, London, UK, September 21-23, 2016. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5090-3506-9. ISSN 1550-6525. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DS-RT.2016.18\" target=\"publication\">10.1109\/DS-RT.2016.18<\/a>. Acceptance rate 42.0% (21\/50). Best paper candidate. <a href=\"javascript:showAbstract('The path to extreme scale high-performance computing (HPC) poses several challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Investigating the performance of parallel applications at scale on future architectures and the performance impact of different architectural choices is an important component of HPC hardware\/software co-design. Simulations using models of future HPC systems and communication traces from applications running on existing HPC systems can offer an insight into the performance of future architectures. This work targets technology developed for scalable application tracing of communication events. It focuses on extreme-scale simulation of HPC applications and their communication behavior via lightweight parallel discrete event simulation for performance estimation and evaluation. Instead of simply replaying a trace within a simulator, this work promotes the generation of a benchmark from traces. This benchmark is subsequently exposed to simulation using models to reflect the performance characteristics of future-generation HPC systems. This technique provides a number of benefits, such as eliminating the data intensive trace replay and enabling simulations at different scales. The presented work features novel software co-design aspects, combining the ScalaTrace tool to generate scalable trace files, the ScalaBenchGen tool to generate the benchmark, and the xSim tool to assess the benchmark characteristics within a simulator.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/lagadapati16benchmark.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/lagadapati16benchmark.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#lagadapati16benchmark\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Havens: Explicit Reliable Memory Regions for HPC Applications<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ieee-hpec.org\" target=\"ieee-hpec.org\">20th IEEE High Performance Extreme Computing Conference (HPEC) 2016<\/a><\/i>, pages 1-6, Waltham, MA, USA, September 13-15, 2016. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/HPEC.2016.7761593\" target=\"publication\">10.1109\/HPEC.2016.7761593<\/a>. <a href=\"javascript:showAbstract('Supporting error resilience in future exascale-class supercomputing systems is a critical challenge. Due to transistor scaling trends and increasing memory density, the scientific simulations are expected to experience more interruptions caused by soft errors in the system memory. Existing hardware-based detection and recovery techniques will be inadequate in the presence of high memory fault rates. In this paper we propose a partial memory protection scheme using region-based memory management. We define regions called havens that provide fault protection for program objects. We provide reliability for the regions through a software-based parity protection mechanism. Our approach enables critical application code and variables to be placed in these havens. The fault coverage of our approach is application agnostic unlike algorithm-based fault tolerance techniques.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar16havens.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar16havens.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar16havens\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Kun Tang, Devesh Tiwari, Saurabh Gupta, Ping Huang, QiQi Lu, Christian Engelmann, and Xubin He. <b>Power-Capping Aware Checkpointing: On the Interplay Among Power-Capping, Temperature, Reliability, Performance, and Energy<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.dsn.org\" target=\"www.dsn.org\">46th IEEE\/IFIP International Conference on Dependable  Systems and Networks (DSN) 2016<\/a><\/i>, pages 311-322, Toulouse, France, June 28 &#8211; July 1, 2016. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISSN 2158-3927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DSN.2016.36\" target=\"publication\">10.1109\/DSN.2016.36<\/a>. Acceptance rate 22.4% (58\/259). <a href=\"javascript:showAbstract('Checkpoint and restart mechanisms have been widely used in large scientific simulation applications to make forward progress in case of failures. However, none of the prior works have considered the interaction of power-constraint with temperature, reliability, performance, and checkpointing interval. It is not clear how power-capping may affect optimal checkpointing interval. What are the involved reliability, performance, and energy trade-offs? In this paper, we develop a deep understanding about the interaction between power-capping and scientific applications using checkpoint\/restart as resilience mechanism, and propose a new model for the optimal checkpointing interval (OCI) under power-capping. Our study reveals several interesting, and previously unknown, insights about how power-capping affects the reliability, energy consumption, performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tang16power-capping.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#tang16power-capping\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Frank Mueller, Kurt Ferreira, and Christian Engelmann. <b>Mini-Ckpts: Surviving OS Failures in Persistent Memory<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ics16.bilkent.edu.tr\" target=\"ics16.bilkent.edu.tr\">30th ACM International Conference on Supercomputing  (ICS) 2016<\/a><\/i>, pages 7:1-7:14, Istanbul, Turkey, June 1-3, 2016. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-4361-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/2925426.2926295\" target=\"publication\">10.1145\/2925426.2926295<\/a>. Acceptance rate 24.2% (43\/178). <a href=\"javascript:showAbstract('Concern is growing in the high-performance computing (HPC) community on the reliability of future extreme-scale systems. Current efforts have focused on application fault-tolerance rather than the operating system (OS), despite the fact that recent studies have suggested that failures in OS memory are more likely. The OS is critical to a system&amp;#39;s correct and efficient operation of the node and processes it governs -- and in HPC also for any other nodes a parallelized application runs on and communicates with: Any single node failure generally forces all processes of this application to terminate due to tight communication in HPC. Therefore, the OS itself must be capable of tolerating failures. In this work, we introduce mini-ckpts, a framework which enables application survival despite the occurrence of a fatal OS failure or crash. Mini-ckpts achieves this tolerance by ensuring that the critical data describing a process is preserved in persistent memory prior to the failure. Following the failure, the OS is rejuvenated via a warm reboot and the application continues execution effectively making the failure and restart transparent. The mini-ckpts rejuvenation and recovery process is measured to take between three to six seconds and has a failure-free overhead of between 3-5% for a number of key HPC workloads. In contrast to current fault-tolerance methods, this work ensures that the operating and runtime system can continue in the presence of faults. This is a much finer-grained and dynamic method of fault-tolerance than the current, coarse-grained, application-centric methods. Handling faults at this level has the potential to greatly reduce overheads and enables mitigation of additional fault scenarios.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala16mini-ckpts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/fiala16mini-ckpts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#fiala16mini-ckpts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Leonardo Bautista-Gomez, Ana Gainaru, Swann Perarnau, Devesh Tiwari, Saurabh Gupta, Franck Cappello, Christian Engelmann, and Marc Snir. <b>Reducing Waste in Extreme Scale Systems Through Introspective Analysis<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ipdps.org\" target=\"www.ipdps.org\">30th IEEE International Parallel and Distributed Processing Symposium (IPDPS) 2016<\/a><\/i>, pages 212-221, Chicago, IL, USA, May 23-27, 2016. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISSN 1530-2075. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/IPDPS.2016.100\" target=\"publication\">10.1109\/IPDPS.2016.100<\/a>. Acceptance rate 23.0% (114\/496). <a href=\"javascript:showAbstract('Resilience is an important challenge for extreme-scale  supercomputers. Today, failures in supercomputers are  assumed to be uniformly distributed in time. However, recent  studies show that failures in high-performance computing  systems are partially correlated in time, generating periods  of higher failure density. Our study of the failure logs of  multiple supercomputers show that periods of higher failure  density occur with up to three times more than the average.  We design a monitoring system that listens to hardware  events and forwards important events to the runtime to  detect those regime changes. We implement a runtime capable  of receiving notifications and adapt dynamically. In  addition, we build an analytical model to predict the gains  that such dynamic approach could achieve. We demonstrate that  in some systems, our approach can reduce the wasted time.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/bautista-gomez16reducing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/bautista-gomez16reducing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#bautista-gomez16reducing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>Supporting the Development of Soft-Error Resilient Message Passing Applications using Simulation<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-795.html\" target=\"www.iasted.org\/conferences\/home-795.html\">13th IASTED International Conference on Parallel and Distributed Computing and Networks (PDCN) 2016<\/a><\/i>, Innsbruck, Austria, February 15-16, 2016. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-979-0. DOI <a href=\"http:\/\/dx.doi.org\/10.2316\/P.2016.834-005\" target=\"publication\">10.2316\/P.2016.834-005<\/a>. <a href=\"javascript:showAbstract('Radiation-induced bit flip faults are of particular concern in extreme-scale high-performance computing systems. This paper presents a simulation-based tool that enables the development of soft-error resilient message passing applications by permitting the investigation of their correctness and performance under various fault conditions. The documented extensions to the Extreme-scale Simulator (xSim) enable the injection of bit flip faults at specific of injection location(s) and fault activation time(s), while supporting a significant degree of configurability of the fault type. Experiments show that the simulation overhead with the new feature is ~2,325% for serial execution and ~1,730% at 128 MPI processes, both with very fine-grain fault injection. Fault injection experiments demonstrate the usefulness of the new feature by injecting bit flips in the input and output matrices of a matrix-matrix multiply application, revealing vulnerability of data structures, masking and error propagation. xSim is the very first simulation-based MPI performance tool that supports both, the injection of process failures and bit flip faults.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16supporting.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann16supporting.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann16supporting\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Amogh Katti, Giuseppe Di Fatta, Thomas Naughton, and Christian Engelmann. <b>Scalable and Fault Tolerant Failure Detection and Consensus<\/b>. In <i>Proceedings of the <a href=\"http:\/\/eurompi2015.bordeaux.inria.fr\" target=\"eurompi2015.bordeaux.inria.fr\">22nd European MPI Users` Group Meeting (EuroMPI) 2015<\/a><\/i>, pages 13:1-13:9, Bordeaux, France, September 21-24, 2015. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-3795-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/2802658.2802660\" target=\"publication\">10.1145\/2802658.2802660<\/a>. Acceptance rate 48.3% (14\/29). <a href=\"javascript:showAbstract('Future extreme-scale high-performance computing systems will be required to work under frequent component failures. The MPI Forum&amp;#39;s User Level Failure Mitigation proposal has introduced an operation (MPI_Comm_shrink) to synchronize the alive processes on the list of failed processes, so that applications can continue to execute even in the presence of failures by adopting algorithm-based fault tolerance techniques. The MPI_Comm_shrink operation requires a fault tolerant failure detection and consensus algorithm. This paper presents and compares two novel failure detection and consensus algorithms to support this operation. The proposed algorithms are based on Gossip protocols and are inherently fault-tolerant and scalable. The proposed algorithms were implemented and tested using the Extreme-scale Simulator. The results show that in both algorithms the number of Gossip cycles to achieve global consensus scales logarithmically with system size. The second algorithm also shows better scalability in terms of memory usage and network bandwidth costs and a perfect synchronization in achieving global consensus.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/katti15scalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/katti15scalable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#katti15scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A Network Contention Model for the Extreme-scale Simulator<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-826.html\" target=\"www.iasted.org\/conferences\/home-826.html\">34th IASTED International Conference on Modelling, Identification and Control (MIC) 2015<\/a><\/i>, Innsbruck, Austria, February 17-18, 2015. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-975-2. DOI <a href=\"http:\/\/dx.doi.org\/10.2316\/P.2015.826-043\" target=\"publication\">10.2316\/P.2015.826-043<\/a>. <a href=\"javascript:showAbstract('The Extreme-scale Simulator (xSim) is a performance investigation toolkit for high-performance computing (HPC) hardware\/software co-design. It permits running a HPC application with millions of concurrent execution threads, while observing its performance in a simulated extreme-scale system. This paper details a newly developed network modeling feature for xSim, eliminating the shortcomings of the existing network modeling capabilities. The approach takes a different path for implementing network contention and bandwidth capacity modeling using a less synchronous and accurate enough model design. With the new network modeling feature, xSim is able to simulate on-chip and on-node networks with reasonable accuracy and overheads.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann15network.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann15network.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann15network\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>Improving the Performance of the Extreme-scale Simulator<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ds-rt.com\/2014\" target=\"ds-rt.com\/2014\">18th IEEE\/ACM International Symposium on Distributed Simulation and Real Time Applications (DS-RT) 2014<\/a><\/i>, pages 198-207, Toulouse, France, October 1-3, 2014. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4799-6143-6. ISSN 1550-6525. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/DS-RT.2014.32\" target=\"publication\">10.1109\/DS-RT.2014.32<\/a>. Best paper candidate. <a href=\"javascript:showAbstract('Investigating the performance of parallel applications at scale on future high-performance computing (HPC) architectures and the performance impact of different architecture choices is an important component of HPC hardware\/software co-design. The Extreme-scale Simulator (xSim) is a simulation-based toolkit for investigating the performance of parallel applications at scale. xSim scales to millions of simulated Message Passing Interface (MPI) processes. The overhead introduced by a simulation tool is an important performance and productivity aspect. This paper documents two improvements to xSim: (1) a new deadlock resolution protocol to reduce the parallel discrete event simulation management overhead and (2) a new simulated MPI message matching algorithm to reduce the oversubscription management overhead. The results clearly show a significant performance improvement, such as by reducing the simulation overhead for running the NAS Parallel Benchmark suite inside the simulator  from 1,020% to 238% for the conjugate gradient (CG) benchmark and from 102% to 0% for the embarrassingly parallel (EP) and benchmark, as well as, from 37,511% to 13,808% for CG and from 3,332% to 204% for EP with accurate process failure simulation.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann14improving.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann14improving.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann14improving\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Christian Engelmann, Geoffroy Vall&eacute;e, and Swen B&ouml;hm. <b>Supporting the Development of Resilient Message Passing Applications using Simulation<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2014.org\" target=\"www.pdp2014.org\">22nd Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2014<\/a><\/i>, pages 271-278, Turin, Italy, February 12-14, 2014. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISSN 1066-6192. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP.2014.74\" target=\"publication\">10.1109\/PDP.2014.74<\/a>. Acceptance rate 32.6% (73\/224). <a href=\"javascript:showAbstract('An emerging aspect of high-performance computing (HPC) hardware\/software co-design is investigating performance under failure. The work in this paper extends the Extreme-scale Simulator (xSim), which was designed for evaluating the performance of message passing interface (MPI) applications on future HPC architectures, with fault-tolerant MPI extensions proposed by the MPI Fault Tolerance Working Group. xSim permits running MPI applications with millions of concurrent MPI ranks, while observing application performance in a simulated extreme-scale system using a lightweight parallel discrete event simulation. The newly added features offer user-level failure mitigation (ULFM) extensions at the simulated MPI layer to support algorithm-based fault tolerance (ABFT). The presented solution permits investigating performance under failure and failure handling of ABFT solutions. The newly enhanced xSim is the very first performance tool that supports ULFM and ABFT.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton14supporting.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton14supporting.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton14supporting\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy Vall&eacute;e, Thomas Naughton, Swen B&ouml;hm, and Christian Engelmann. <b>A Runtime Environment for Supporting Research in Resilient HPC System Software &#038; Tools<\/b>. In <i>Proceedings of the <a href=\"http:\/\/is-candar.org\" target=\"is-candar.org\">1st International Symposium on Computing and Networking &#8211; Across Practical Development and Theoretical Research &#8211; (CANDAR) 2013<\/a><\/i>, pages 213-219, Matsuyama, Japan, December 4-6, 2013. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4799-2795-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CANDAR.2013.38\" target=\"publication\">10.1109\/CANDAR.2013.38<\/a>. Acceptance rate 35.8% (28\/78). <a href=\"javascript:showAbstract('The high-performance computing (HPC) community continues to increase the size and complexity of hardware platforms that support advanced scientific workloads. The runtime environment (RTE) is a crucial layer in the software stack for these large-scale systems. The RTE manages the interface between the operating system and the application running in parallel on the machine. The deployment of applications and tools on large-scale HPC computing systems requires the RTE to manage process creation in a scalable manner, support sparse connectivity, and provide fault tolerance. We have developed a new RTE that provides a basis for building distributed execution environments and developing tools for HPC to aid research in system software and resilience. This paper describes the software architecture of the Scalable runTime Component Infrastructure (STCI), which is intended to provide a complete infrastructure for scalable start-up and management of many processes in large-scale HPC systems. We highlight features of the current implementation, which is provided as a system library that allows developers to easily use and integrate STCI in their tools and\/or applications. The motivation for this work has been to support ongoing research activities in fault-tolerance for large-scale systems. We discuss the advantages of the modular framework employed and describe two use cases that demonstrate its capabilities: (i) an alternate runtime for a Message Passing Interface (MPI) stack, and (ii) a distributed control and communication substrate for a fault-injection tool.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/vallee13runtime.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/vallee13runtime.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#vallee13runtime\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Investigating Operating System Noise in Extreme-Scale High-Performance Computing Systems using Simulation<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-795.html\" target=\"www.iasted.org\/conferences\/home-795.html\">11th IASTED International Conference on Parallel and Distributed Computing and Networks (PDCN) 2013<\/a><\/i>, Innsbruck, Austria, February 11-13, 2013. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-943-1. DOI <a href=\"http:\/\/dx.doi.org\/10.2316\/P.2013.795-010\" target=\"publication\">10.2316\/P.2013.795-010<\/a>. <a href=\"javascript:showAbstract('Hardware\/software co-design for future-generation high-performance computing (HPC) systems aims at closing the gap between the peak capabilities of the hardware and the performance realized by applications (application-architecture performance gap). Performance profiling of architectures and applications is a crucial part of this iterative process. The work in this paper focuses on operating system (OS) noise as an additional factor to be considered for co-design. It represents the first step in including OS noise in HPC hardware\/software co-design by adding a noise injection feature to an existing simulation-based co-design toolkit. It reuses an existing abstraction for OS noise with frequency (periodic recurrence) and period (duration of each occurrence) to enhance the processor model of the Extreme-scale Simulator (xSim) with synchronized and random OS noise simulation. The results demonstrate this capability by evaluating the impact of OS noise on MPI_Bcast() and MPI_Reduce() in a simulated future-generation HPC system with 2,097,152 compute nodes.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13investigating.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann13investigating.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann13investigating\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Frank Mueller, Christian Engelmann, Kurt Ferreira, Ron Brightwell, and Rolf Riesen. <b>Detection and Correction of Silent Data Corruption for Large-Scale High-Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc12.supercomputing.org\" target=\"sc12.supercomputing.org\">25th IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2012<\/a><\/i>, pages 78:1-78:12, Salt Lake City, UT, USA, November 10-16, 2012. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4673-0804-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SC.2012.49\" target=\"publication\">10.1109\/SC.2012.49<\/a>. Acceptance rate 21.2% (100\/472). <a href=\"javascript:showAbstract('Faults have become the norm rather than the exception for high-end computing on clusters with 10s\/100s of thousands of cores. Exacerbating this situation, some of these faults remain undetected, manifesting themselves as silent errors that corrupt memory while applications continue to operate and report incorrect results. This paper studies the potential for redundancy to both detect and correct soft errors in MPI message-passing applications. Our study investigates the challenges inherent to detecting soft errors within MPI application while providing transparent MPI redundancy. By assuming a model wherein corruption in application data manifests itself by producing differing MPI message data between replicas, we study the best suited protocols for detecting and correcting MPI data that is the result of corruption. To experimentally validate our proposed detection and correction protocols, we introduce RedMPI, an MPI library which resides in the MPI profiling layer. RedMPI is capable of both online detection and correction of soft errors that occur in MPI applications without requiring any modifications to the application source by utilizing either double or triple redundancy. Our results indicate that our most efficient consistency protocol can successfully protect applications experiencing even high rates of silent data corruption with runtime overheads between 0% and 30% as compared to unprotected applications without redundancy. Using our fault injector within RedMPI, we observe that even a single soft error can have profound effects on running applications, causing a cascading pattern of corruption in most cases causes that spreads to all other processes. RedMPI&amp;#39;s protection has been shown to successfully mitigate the effects of soft errors while allowing applications to complete with correct results even in the face of errors.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala12detection2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/fiala12detection2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#fiala12detection2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>James Elliott, Kishor Kharbas, David Fiala, Frank Mueller, Kurt Ferreira, and Christian Engelmann. <b>Combining Partial Redundancy and Checkpointing for HPC<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icdcs-2012.org\/\" target=\"icdcs-2012.org\/\">32nd International Conference on Distributed Computing Systems (ICDCS) 2012<\/a><\/i>, pages 615-626, Macau, SAR, China, June 18-21, 2012. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4685-8. ISSN 1063-6927. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICDCS.2012.56\" target=\"publication\">10.1109\/ICDCS.2012.56<\/a>. Acceptance rate 13.8% (71\/515). <a href=\"javascript:showAbstract('Today&amp;#39;s largest High Performance Computing (HPC) systems exceed one Petaflops (10^15 floating point operations per second) and exascale systems are projected within seven years. But reliability is becoming one of the major challenges faced by exascale computing. With billion-core parallelism, the mean time to failure is projected to be in the range of minutes or hours instead of days. Failures are becoming the norm rather than the exception during execution of HPC applications. Current fault tolerance techniques in HPC focus on reactive ways to mitigate faults, namely via checkpoint and restart (C\/R). Apart from storage overheads, C\/R-based fault recovery comes at an additional cost in terms of application performance because normal execution is disrupted when checkpoints are taken. Studies have shown that applications running at a large scale spend more than 50% of their total time saving checkpoints, restarting and redoing lost work. Redundancy is another fault tolerance technique, which employs redundant processes performing the same task. If a process fails, a replica of it can take over its execution. Thus, redundant copies can decrease the overall failure rate. The downside of redundancy is that extra resources are required and there is an additional overhead on communication and synchronization. This work contributes a model and analyzes the benefit of C\/R in coordination with redundancy at different degrees to minimize the total wallclock time and resources utilization of HPC applications. We further conduct experiments with an implementation of redundancy within the MPI layer on a cluster. Our experimental results confirm the benefit of dual and triple redundancy - but not for partial redundancy - and show a close fit to the model. At 80,000 processes, dual redundancy requires twice the number of processing resources for an application but allows two jobs of 128 hours wallclock time to finish within the time of just one job without redundancy. For narrow ranges of processor counts, partial redundancy results in the lowest time. Once the count exceeds 770, 000, triple redundancy has the lowest overall cost. Thus, redundancy allows one to trade-off additional resource requirements against wallclock time, which provides a tuning knob for users to adapt to resource availabilities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/elliott12combining.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/elliott12combining.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#elliott12combining\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Sudharshan S. Vazhkudai, Xiaosong Ma, Fei Meng, Youngjae Kim, and Christian Engelmann. <b>NVMalloc: Exposing an Aggregate SSD Store as a Memory Partition in Extreme-Scale Machines<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ipdps.org\" target=\"www.ipdps.org\">26th IEEE International Parallel and Distributed Processing Symposium (IPDPS) 2012<\/a><\/i>, pages 957-968, Shanghai, China, May 21-25, 2012. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4675-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/IPDPS.2012.90\" target=\"publication\">10.1109\/IPDPS.2012.90<\/a>. Acceptance rate 20.7% (118\/569). <a href=\"javascript:showAbstract('DRAM is a precious resource in extreme-scale machines and is increasingly becoming scarce, mainly due to the growing number of cores per node. On future multi-petaflop and exaflop machines, the memory pressure is likely to be so severe that we need to rethink our memory usage models. Fortunately, the advent of non-volatile memory (NVM) offers a unique opportunity in this space. Current NVM offerings possess several desirable properties, such as low cost and power efficiency, but also suffer from high latency and lifetime issues. We need rich techniques to be able to use them alongside DRAM. In this paper, we propose a novel approach to exploiting NVM as a secondary memory partition so that applications can explicitly allocate and manipulate memory regions therein. More specifically, we propose an NVMalloc library with a suite of services that enables applications to access a distributed NVM storage system. We have devised ways within NVMalloc so that the storage system, built from compute node-local NVM devices, can be accessed in a byte-addressable fashion using the memory mapped I\/O interface. Our approach has the potential to re-energize out-of-core computations on large-scale machines by having applications allocate certain variables through NVMalloc, thereby increasing the overall memory available for the application. Our evaluation on a 128-core cluster shows that NVMalloc enables applications to compute problem sizes larger than the physical memory in a cost-effective manner. It can achieve better performance with increased computation time between NVM memory accesses or increased data access locality. In addition, our results suggest that while NVMalloc enables transparent access to NVM-resident variables, the explicit control it provides is crucial to optimize application performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang12nvmalloc.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/wang12nvmalloc.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#wang12nvmalloc\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen B&ouml;hm and Christian Engelmann. <b>File I\/O for MPI Applications in Redundant Execution Scenarios<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2012.org\" target=\"www.pdp2012.org\">20th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2012<\/a><\/i>, pages 112-119, Garching, Germany, February 15-17, 2012. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4633-9. ISSN 1066-6192. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP.2012.22\" target=\"publication\">10.1109\/PDP.2012.22<\/a>. <a href=\"javascript:showAbstract('As multi-petascale and exa-scale high-performance computing (HPC) systems inevitably have to deal with a number of resilience challenges, such as a significant growth in component count and smaller circuit sizes with lower circuit voltages, redundancy may offer an acceptable level of resilience that traditional fault tolerance techniques, such as checkpoint\/restart, do not. Although redundancy in HPC is quite controversial due to the associated cost for redundant components,  the constantly increasing number of cores-per-processor is tilting this cost calculation toward a system design where computation, such as for redundancy, is much cheaper and communication, needed for checkpoint\/restart, is much more expensive. Recent research and development activities in redundancy for Message Passing Interface (MPI) applications focused on availability\/reliability models and replication algorithms. This paper takes a first step toward solving an open research problem associated with running a parallel application redundantly, which is file I\/O under redundancy. The approach intercepts file I\/O calls made by a redundant application to employ coordination protocols that execute file I\/O operations in a redundancy-oblivious fashion when accessing a node-local file system, or in a redundancy-aware fashion when accessing a shared networked file system. A proof-of concept prototype is presented and a number of coordination protocols are described and evaluated. The results show the performance impact for redundantly accessing a shared networked file system, but also demonstrate the capability to regain performance by utilizing MPI communication between replicas and parallel file I\/O.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm12file.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/boehm12file.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#boehm12file\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen B&ouml;hm and Christian Engelmann. <b>xSim: The Extreme-Scale Simulator<\/b>. In <i>Proceedings of the <a href=\"http:\/\/hpcs11.cisedu.info\" target=\"hpcs11.cisedu.info\">International Conference on High Performance Computing and Simulation (HPCS) 2011<\/a><\/i>, pages 280-286, Istanbul, Turkey, July 4-8, 2011. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-61284-383-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/HPCSim.2011.5999835\" target=\"publication\">10.1109\/HPCSim.2011.5999835<\/a>. Acceptance rate 28.1% (48\/171). <a href=\"javascript:showAbstract('Investigating parallel application performance properties at scale is becoming an important part of high-performance computing (HPC) application development and deployment. The Extreme-scale Simulator (xSim) is a performance investigation toolkit that permits running an application in a controlled environment at extreme scale without the need for a respective extreme-scale HPC system. Using a lightweight parallel discrete event simulation, xSim executes a parallel application with a virtual wall clock time, such that performance data can be extracted based on a processor model and a network model. This paper presents significant enhancements to the xSim toolkit prototype that provide a more complete Message Passing Interface (MPI) support and improve its versatility. These enhancements include full virtual MPI group, communicator and collective communication support, and global variables support. The new capabilities are demonstrated by executing the entire NAS Parallel Benchmark suite in a simulated HPC environment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm11xsim.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/boehm11xsim.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#boehm11xsim\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Swen B&ouml;hm. <b>Redundant Execution of HPC Applications with MR-MPI<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-719.html\" target=\"www.iasted.org\/conferences\/home-719.html\">10th IASTED International Conference on Parallel and Distributed Computing and Networks (PDCN) 2011<\/a><\/i>, pages 31-38, Innsbruck, Austria, February 15-17, 2011. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-864-9. DOI <a href=\"http:\/\/dx.doi.org\/10.2316\/P.2011.719-031\" target=\"publication\">10.2316\/P.2011.719-031<\/a>. <a href=\"javascript:showAbstract('This paper presents a modular-redundant Message Passing Interface (MPI) solution, MR-MPI, for transparently executing  high-performance computing (HPC) applications in a redundant fashion. The presented work addresses the deficiencies of recovery-oriented HPC, i.e., checkpoint\/restart to\/from a parallel file system, at extreme scale by adding the redundancy approach to the HPC resilience portfolio. It utilizes the MPI performance tool interface, PMPI, to transparently intercept MPI calls from an application and to hide all redundancy-related mechanisms. A redundantly executed application runs with &amp;#36;r*m native MPI processes, where r is the number of MPI ranks visible to the application and m is the replication degree. Messages between redundant nodes are replicated. Partial replication for tunable resilience is supported. The performance results clearly show the negative impact of the O(m^2) messages between replicas. For low-level, point-to-point benchmarks, the impact can be as high as the replication degree. For applications, performance highly depends on the actual communication types and counts. On single-core systems, the overhead can be 0% for embarrassingly parallel applications independent of the employed redundancy configuration or up to 70-90% for communication-intensive applications in a dual-redundant configuration. On multi-core systems, the overhead can be significantly higher due to the additional communication contention.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann11redundant.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann11redundant.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann11redundant\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Hybrid Checkpointing for MPI Jobs in HPC Environments<\/b>. In <i>Proceedings of the <a href=\"http:\/\/grid.sjtu.edu.cn\/icpads10\" target=\"grid.sjtu.edu.cn\/icpads10\">16th IEEE International Conference on Parallel and Distributed Systems (ICPADS) 2010<\/a><\/i>, pages 524-533, Shanghai, China, December 8-10, 2010. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4307-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICPADS.2010.48\" target=\"publication\">10.1109\/ICPADS.2010.48<\/a>. Acceptance rate 29.6% (77\/188). <a href=\"javascript:showAbstract('As the core count in high-performance computing systems keeps increasing, faults are becoming common place. Check pointing addresses such faults but captures full process images even though only a subset of the process image changes between checkpoints. We have designed a hybrid check pointing technique for MPI tasks of high-performance applications. This technique alternates between full and incremental checkpoints: At incremental checkpoints, only data changed since the last checkpoint is captured. Our implementation integrates new BLCR and LAM\/MPI features that complement traditional full checkpoints. This results in significantly reduced checkpoint sizes and overheads with only moderate increases in restart overhead. After accounting for cost and savings, benefits due to incremental checkpoints are an order of magnitude larger than overheads on restarts. We further derive qualitative results indicating an optimal balance between full\/incremental checkpoints of our novel approach at a ratio of 1:9, which outperforms both always-full and always-incremental check pointing.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang10hybrid2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/wang10hybrid2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#wang10hybrid2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Min Li, Sudharshan S. Vazhkudai, Ali R. Butt, Fei Meng, Xiaosong Ma, Youngjae Kim, Christian Engelmann, and Galen Shipman. <b>Functional Partitioning to Optimize End-to-End Performance on Many-Core Architectures<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc10.supercomputing.org\" target=\"sc10.supercomputing.org\">23rd IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2010<\/a><\/i>, pages 1-12, New Orleans, LA, USA, November 13-19, 2010. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4244-7559-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SC.2010.28\" target=\"publication\">10.1109\/SC.2010.28<\/a>. Acceptance rate 19.8% (50\/253). <a href=\"javascript:showAbstract('Scaling computations on emerging massive-core supercomputers is a daunting task, which coupled with the significantly lagging system I\/O capabilities exacerbates applications&amp;#39; end-to-end performance. The I\/O bottleneck often negates potential performance benefits of assigning additional compute cores to an application. In this paper, we address this issue via a novel functional partitioning (FP) runtime environment that allocates cores to specific application tasks - checkpointing, de-duplication, and scientific data format transformation - so that the deluge of cores can be brought to bear on the entire gamut of application activities. The focus is on utilizing the extra cores to support HPC application I\/O activities and also leverage solid-state disks in this context. For example, our evaluation shows that dedicating 1 core on an oct-core machine for checkpointing and its assist tasks using FP can improve overall execution time of a FLASH benchmark on 80 and  160 cores by 43.95% and 41.34%, respectively.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/li10functional.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/li10functional.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#li10functional\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen B&ouml;hm, Christian Engelmann, and Stephen L. Scott. <b>Aggregation of Real-Time System Monitoring Data for Analyzing Large-Scale Parallel and Distributed Computing Environments<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.anss.org.au\/hpcc2010\" target=\"www.anss.org.au\/hpcc2010\">12th IEEE International Conference on High Performance Computing and Communications (HPCC) 2010<\/a><\/i>, pages 72-78, Melbourne, Australia, September 1-3, 2010. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4214-0. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/HPCC.2010.32\" target=\"publication\">10.1109\/HPCC.2010.32<\/a>. Acceptance rate 19.1% (58\/304). <a href=\"javascript:showAbstract('We present a monitoring system for large-scale parallel and distributed computing environments that allows to trade-off accuracy in a tunable fashion to gain scalability without compromising fidelity. The approach relies on classifying each gathered monitoring metric based on individual needs and on aggregating messages containing classes of individual monitoring metrics using a tree-based overlay network. The MRNet-based prototype is able to significantly reduce the amount of gathered and stored monitoring data, e.g., by a factor of  56 in comparison to the Ganglia distributed monitoring system. A simple scaling study reveals, however, that further efforts are needed in reducing the amount of data to monitor future-generation extreme-scale systems with up to 1,000,000 nodes. The implemented solution did not had a measurable performance impact as the 32-node test system did not produce enough monitoring data to interfere with running applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm10aggregation.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/boehm10aggregation.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#boehm10aggregation\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Antonina Litvinova, Christian Engelmann, and Stephen L. Scott. <b>A Proactive Fault Tolerance Framework for High-Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-676.html\" target=\"www.iasted.org\/conferences\/home-676.html\">9th IASTED International Conference on Parallel and Distributed Computing and Networks (PDCN) 2010<\/a><\/i>, Innsbruck, Austria, February 16-18, 2010. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-783-3. DOI <a href=\"http:\/\/dx.doi.org\/10.2316\/P.2010.676-024\" target=\"publication\">10.2316\/P.2010.676-024<\/a>. <a href=\"javascript:showAbstract('As high-performance computing (HPC) systems continue to increase in scale, their mean-time to interrupt decreases respectively. The current state of practice for fault tolerance (FT) is checkpoint\/restart. However, with increasing error rates, increasing aggregate memory and not proportionally increasing I\/O capabilities, it is becoming less efficient. Proactive FT avoids experiencing failures through preventative measures, such as by migrating application parts away from nodes that are about to fail. This paper presents a proactive FT framework that performs environmental monitoring, event logging, parallel job monitoring and resource monitoring to analyze HPC system reliability and to perform FT through such preventative actions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/litvinova10proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/litvinova10proactive.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#litvinova10proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Narate Taerat, Nichamon Naksinehaboon, Clayton Chandler, James Elliott, Chokchai (Box) Leangsuksun, George Ostrouchov, Stephen L. Scott, and Christian Engelmann. <b>Blue Gene\/L Log Analysis and Time to Interrupt Estimation<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ares-conference.eu\/ares2009\" target=\"www.ares-conference.eu\/ares2009\">4th International Conference on Availability, Reliability and Security (ARES) 2009<\/a><\/i>, pages 173-180, Fukuoka, Japan, March 16-19, 2009. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4244-3572-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ARES.2009.105\" target=\"publication\">10.1109\/ARES.2009.105<\/a>. Acceptance rate 25.0% (40\/160). <a href=\"javascript:showAbstract('System- and application-level failures could be characterized by analyzing relevant log files. The resulting data might then be used in numerous studies on and future developments for the mission-critical and large scale computational architecture, including fields such as failure prediction, reliability modeling, performance modeling and power awareness. In this paper, system logs covering a six month period of the Blue Gene\/L supercomputer were obtained and subsequently analyzed. Temporal filtering was applied to remove duplicated log messages. Optimistic and pessimistic perspectives were exerted on filtered log information to observe failure behavior within the system. Further, various time to repair factors were applied to obtain application time to interrupt, which will be exploited in further resilience modeling research.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/taerat09blue.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#taerat09blue\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Hong H. Ong, and Stephen L. Scott. <b>Evaluating the Shared Root File System Approach for Diskless High-Performance Computing Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.linuxclustersinstitute.org\/conferences\" target=\"www.linuxclustersinstitute.org\/conferences\">10th LCI International Conference on High-Performance Clustered Computing (LCI) 2009<\/a><\/i>, Boulder, CO, USA, March 9-12, 2009. <a href=\"javascript:showAbstract('Diskless high-performance computing (HPC) systems utilizing networked storage have become popular in the last several years. Removing disk drives significantly increases compute node reliability as they are known to be a major source of failures. Furthermore, networked storage solutions utilizing parallel I\/O and replication are able to provide increased scalability and availability. Reducing a compute node to processor(s), memory and network interface(s) greatly reduces its physical size, which in turn allows for large-scale dense HPC solutions. However, one major obstacle is the requirement by certain operating systems (OSs), such as Linux, for a root file system. While one solution is to remove this requirement from the OS, another is to share the root file system over the networked storage. This paper evaluates three networked file system solutions, NFSv4, Lustre and PVFS2, with respect to their performance, scalability, and availability features for servicing a common root file system in a diskless HPC configuration. Our findings indicate that Lustre is a viable solution as it meets both, scaling and performance requirements. However, certain availability issues regarding single points of failure and control need to be considered.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09evaluating.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann09evaluating.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09evaluating\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, and Stephen L. Scott. <b>Proactive Fault Tolerance Using Preemptive Migration<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2009.org\" target=\"www.pdp2009.org\">17th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2009<\/a><\/i>, pages 252-257, Weimar, Germany, February 18-20, 2009. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3544-9. ISSN 1066-6192. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP.2009.31\" target=\"publication\">10.1109\/PDP.2009.31<\/a>. Acceptance rate 42.0% (58\/138). <a href=\"javascript:showAbstract('Proactive fault tolerance (FT) in high-performance computing is a concept that prevents compute node failures from impacting running parallel applications by preemptively migrating application parts away from nodes that are about to fail. This paper provides a foundation for proactive FT by defining its architecture and classifying implementation options. This paper further relates prior work to the presented architecture and classification, and discusses the challenges ahead for needed supporting technologies.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann09proactive.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Alessandro Valentini, Christian Di Biagio, Fabrizio Batino, Guido Pennella, Fabrizio Palma, and Christian Engelmann. <b>High Performance Computing with Harness over InfiniBand<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2009.org\" target=\"www.pdp2009.org\">17th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2009<\/a><\/i>, pages 151-154, Weimar, Germany, February 18-20, 2009. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3544-9. ISSN 1066-6192. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP.2009.64\" target=\"publication\">10.1109\/PDP.2009.64<\/a>. Acceptance rate 42.0% (58\/138). <a href=\"javascript:showAbstract('Harness is an adaptable and plug-in-based middleware framework able to support distributed parallel computing. By now, it is based on the Ethernet protocol which cannot guarantee high performance throughput and Real Time (determinism) performance. During last years, both the research and industry environments have developed both new network architectures (InfiniBand, Myrinet, iWARP, etc.) to avoid those limits. This paper concerns the integration between Harness and InfiniBand focusing on two solutions: IP over InfiniBand (IPoIB) and Socket Direct Protocol (SDP) technology. Those allow Harness middleware to take advantage of the enhanced features provided by InfiniBand.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/valentini09high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#valentini09high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Hong H. Ong, and Stephen L. Scott. <b>The Case for Modular Redundancy in Large-Scale High Performance Computing Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-641.html\" target=\"www.iasted.org\/conferences\/home-641.html\">8th IASTED International Conference on Parallel and Distributed Computing and Networks (PDCN) 2009<\/a><\/i>, pages 189-194, Innsbruck, Austria, February 16-18, 2009. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-784-0. <a href=\"javascript:showAbstract('Recent investigations into resilience of large-scale high-performance computing (HPC) systems showed a continuous trend of decreasing reliability and availability. Newly installed systems have a lower mean-time to failure (MTTF) and a higher mean-time to recover (MTTR) than their predecessors. Modular redundancy is being used in many mission critical systems today to provide for resilience, such as for aerospace and command &amp; control systems. The primary argument against modular redundancy for resilience in HPC has always been that the capability of a HPC system, and respective return on investment, would be significantly reduced. We argue that modular redundancy can significantly increase compute node availability as it removes the impact of scale from single compute node MTTR. We further argue that single compute nodes can be much less reliable, and therefore less expensive, and still be highly available, if their MTTR\/MTTF ratio is maintained.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09case.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann09case.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09case\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Proactive Process-Level Live Migration in HPC Environments<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc08.supercomputing.org\" target=\"sc08.supercomputing.org\">21st IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2008<\/a><\/i>, pages 1-12, Austin, TX, USA, November 15-21, 2008. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4244-2835-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1413370.1413414\" target=\"publication\">10.1145\/1413370.1413414<\/a>. Acceptance rate 21.3% (59\/277). <a href=\"javascript:showAbstract('As the number of nodes in high-performance computing environments keeps increasing, faults are becoming common place. Reactive fault tolerance (FT) often does not scale due to massive I\/O requirements and relies on manual job resubmission. This work complements reactive with proactive FT at the process level. Through health monitoring, a subset of node failures can be anticipated when one&amp;#39;s health deteriorates. A novel process-level live migration mechanism supports continued execution of applications during much of processes migration. This scheme is integrated into an MPI execution environment to transparently sustain health-inflicted node failures, which eradicates the need to restart and requeue MPI jobs. Experiments indicate that 1-6.5 seconds of prior warning are required to successfully trigger live process migration while similar operating system virtualization mechanisms require 13-24 seconds. This self-healing approach complements reactive FT by nearly cutting the number of checkpoints in half when 70% of the faults are handled proactively.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang08proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/wang08proactive.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#wang08proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Symmetric Active\/Active Replication for Dependent Services<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ares-conference.eu\/ares2008\" target=\"www.ares-conference.eu\/ares2008\">3rd International Conference on Availability, Reliability and Security (ARES) 2008<\/a><\/i>, pages 260-267, Barcelona, Spain, March 4-7, 2008. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3102-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ARES.2008.64\" target=\"publication\">10.1109\/ARES.2008.64<\/a>. Acceptance rate 21.1% (40\/190). <a href=\"javascript:showAbstract('During the last several years, we have established the symmetric active\/active replication model for service-level high availability and implemented several proof-of-concept prototypes. One major deficiency of our model is its inability to deal with dependent services, since its original architecture is based on the client-service model. This paper extends our model to dependent services using its already existing mechanisms and features. The presented concept is based on the idea that a service may also be a client of another service, and multiple services may be clients of each other. A high-level abstraction is used to illustrate dependencies between clients and services, and to decompose dependencies between services into respective client-service dependencies. This abstraction may be used for providing high availability in distributed computing systems with complex service-oriented architectures.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08symmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann08symmetric.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08symmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy R. Vall&eacute;e, Kulathep Charoenpornwattana, Christian Engelmann, Anand Tikotekar, Chokchai (Box) Leangsuksun, Thomas Naughton, and Stephen L. Scott. <b>A Framework For Proactive Fault Tolerance<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ares-conference.eu\/ares2008\" target=\"www.ares-conference.eu\/ares2008\">3rd International Conference on Availability, Reliability and Security (ARES) 2008<\/a><\/i>, pages 659-664, Barcelona, Spain, March 4-7, 2008. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3102-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ARES.2008.171\" target=\"publication\">10.1109\/ARES.2008.171<\/a>. Acceptance rate 21.1% (40\/190). <a href=\"javascript:showAbstract('Fault tolerance is a major concern to guarantee availability of critical services as well as application execution. Traditional approaches for fault tolerance include checkpoint\/restart or duplication. However it is also possible to anticipate failures and proactively take action before failures occur in order to minimize failure impact on the system and application execution. This document presents a proactive fault tolerance framework. This framework can use different proactive fault tolerance mechanisms, i.e. migration and pause\/unpause. The framework also allows the implementation of new proactive fault tolerance policies thanks to a modular architecture. A first proactive fault tolerance policy has been implemented and preliminary experimentations have been done based on system-level virtualization and compared with results obtained by simulation.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/vallee08framework.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/vallee08framework.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#vallee08framework\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Bj&ouml;rn K&ouml;nning, Christian Engelmann, Stephen L. Scott, and George A. (Al) Geist. <b>Virtualized Environments for the Harness High Performance Computing Workbench<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2008.org\" target=\"www.pdp2008.org\">16th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2008<\/a><\/i>, pages 133-140, Toulouse, France, February 13-15, 2008. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3089-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP.2008.14\" target=\"publication\">10.1109\/PDP.2008.14<\/a>. Acceptance rate 40% (83\/207). <a href=\"javascript:showAbstract('This paper describes recent accomplishments in providing a virtualized environment concept and prototype for scientific application development and deployment as part of the Harness High Performance Computing (HPC) Workbench research effort. The presented work focuses on tools and mechanisms that simplify scientific application development and deployment tasks, such that only minimal adaptation is needed when moving from one HPC system to another or after HPC system upgrades. The overall technical approach focuses on the concept of adapting the HPC system environment to the actual needs of individual scientific applications instead of the traditional scheme of adapting scientific applications to individual HPC system environment properties. The presented prototype implementation is based on the mature and lightweight chroot virtualization approach for Unix-type systems with a focus on virtualized file system structure and virtualized shell environment variables utilizing virtualized environment configuration descriptions in Extensible Markup Language (XML) format. The presented work can be easily extended to other virtualization technologies, such as system-level virtualization solutions using hypervisors.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/koenning08virtualized.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/koenning08virtualized.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#koenning08virtualized\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, Hong H. Ong, and Stephen L. Scott. <b>System-level Virtualization for High Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.pdp2008.org\" target=\"www.pdp2008.org\">16th Euromicro International Conference on Parallel, Distributed, and network-based Processing (PDP) 2008<\/a><\/i>, pages 636-643, Toulouse, France, February 13-15, 2008. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3089-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/PDP.2008.85\" target=\"publication\">10.1109\/PDP.2008.85<\/a>. Acceptance rate 40% (83\/207). <a href=\"javascript:showAbstract('System-level virtualization has been a research topic since the 70`s but regained popularity during the past few years because of the availability of efficient solution such as Xen and the implementation of hardware support in commodity processors (e.g. Intel-VT, AMD-V). However, a majority of system-level virtualization projects is guided by the server consolidation market. As a result, current virtualization solutions appear to not be suitable for high performance computing (HPC) which is typically based on large-scale systems. On another hand there is significant interest in exploiting virtual machines (VMs) within HPC for a number of other reasons. By virtualizing the machine, one is able to run a variety of operating systems and environments as needed by the applications. Virtualization allows users to isolate workloads, improving security and reliability. It is also possible to support non-native environments and\/or legacy operating environments through virtualization. In addition, it is possible to balance work loads, use migration techniques to relocate applications from failing machines, and isolate fault systems for repair. This document presents the challenges for the implementation of a system-level virtualization solution for HPC. It also presents a brief survey of the different approaches and techniques to address these challenges.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/vallee08system.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/vallee08system.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#vallee08system\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Li Ou, Christian Engelmann, Xubin (Ben) He, Xin Chen, and Stephen L. Scott. <b>Symmetric Active\/Active Metadata Service for Highly Available Cluster Storage Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.iasted.org\/conferences\/home-590.html\" target=\"www.iasted.org\/conferences\/home-590.html\">19th IASTED International Conference on Parallel and Distributed Computing and Systems (PDCS) 2007<\/a><\/i>, Cambridge, MA, USA, November 19-21, 2007. <a href=\"http:\/\/www.actapress.com\" target=\"www.actapress.com\">ACTA Press, Calgary, AB, Canada<\/a>. ISBN 978-0-88986-703-1. Acceptance rate 49%. <a href=\"javascript:showAbstract('In a typical distributed storage system, metadata is stored and managed by dedicated metadata servers. One way to improve the availability of distributed storage systems is to deploy multiple metadata servers. Past research focused on the active\/standby model, where each active server has at least one redundant idle backup. However, interruption of service and loss of service state may occur during a fail-over depending on the used replication technique. The research in this paper targets the symmetric active\/active replication model using multiple redundant service nodes running in virtual synchrony. In this model, service node failures do not cause a fail-over to a backup and there is no disruption of service or loss of service state. We propose a fast delivery protocol to reduce the latency of total order broadcast. Our prototype implementation shows that high availability of metadata servers can be achieved with an acceptable performance trade-off using the active\/active metadata server solution.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ou07symmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ou07symmetric.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ou07symmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Emanuele Di Saverio, Marco Cesati, Christian Di Biagio, Guido Pennella, and Christian Engelmann. <b>Distributed Real-Time Computing with Harness<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/pvmmpi07.lri.fr\" target=\"pvmmpi07.lri.fr\">14th European PVM\/MPI Users` Group Meeting (EuroPVM\/MPI) 2007<\/a><\/i>, pages 281-288, Paris, France, September 30 &#8211; October 3, 2007. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-540-75415-2. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-540-75416-9_39\" target=\"publication\">10.1007\/978-3-540-75416-9_39<\/a>. <a href=\"javascript:showAbstract('Modern parallel and distributed computing solutions are often built onto a middleware software layer providing a higher and common level of service between computational nodes. Harness is an adaptable, plugin-based middleware framework for parallel and distributed computing. This paper reports recent research and development results of using Harness for real-time distributed computing applications in the context of an industrial environment with the needs to perform several safety critical tasks. The presented work exploits the modular architecture of Harness in conjunction with a lightweight threaded implementation to resolve several real-time issues by adding three new Harness plug-ins to provide a prioritized lightweight execution environment, low latency communication facilities, and local timestamped event logging.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/disaverio07distributed.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/disaverio07distributed.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#disaverio07distributed\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Li Ou, Xubin (Ben) He, Christian Engelmann, and Stephen L. Scott. <b>A Fast Delivery Protocol for Total Order Broadcasting<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.icccn.org\/icccn07\" target=\"www.icccn.org\/icccn07\">16th IEEE International Conference on Computer Communications and Networks (ICCCN) 2007<\/a><\/i>, pages 730-734, Honolulu, HI, USA, August 13-16, 2007. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-42441-251-8. ISSN 1095-2055. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICCCN.2007.4317904\" target=\"publication\">10.1109\/ICCCN.2007.4317904<\/a>. Acceptance rate 29.1% (160\/550). <a href=\"javascript:showAbstract('Sequencer, privilege-based, and communication history algorithms are popular approaches to implement total ordering, where communication history algorithms are most suitable for parallel computing systems, because they provide best performance under heavy work load. Unfortunately, post-transmission delay of communication history algorithms is most apparent when a system is idle. In this paper, we propose a fast delivery protocol to reduce the latency of message ordering. The protocol optimizes the total ordering process by waiting for messages only from a subset of the machines in the group, and by fast acknowledging messages on behalf of other machines. Our test results indicate that the fast delivery protocol is suitable for both idle and heavy load systems, while reducing the latency of message ordering.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ou07fast.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ou07fast.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ou07fast\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Arun B. Nagarajan, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Proactive Fault Tolerance for HPC with Xen Virtualization<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ics07.ac.upc.edu\" target=\"ics07.ac.upc.edu\">21st ACM International Conference on Supercomputing (ICS) 2007<\/a><\/i>, pages 23-32, Seattle, WA, USA, June 16-20, 2007. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-59593-768-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1274971.1274978\" target=\"publication\">10.1145\/1274971.1274978<\/a>. Acceptance rate 23.6% (29\/123). <a href=\"javascript:showAbstract('Large-scale parallel computing is relying increasingly on clusters with thousands of processors. At such large counts of compute nodes, faults are becoming common place. Current techniques to tolerate faults focus on reactive schemes to recover from faults and generally rely on a checkpoint\/restart mechanism. Yet, in today`s systems, node failures can often be anticipated by detecting a deteriorating health status. Instead of a reactive scheme for fault tolerance (FT), we are promoting a proactive one where processes automatically migrate from unhealthy nodes to healthy ones. Our approach relies on operating system virtualization techniques exemplified by but not limited to Xen. This paper contributes an automatic and transparent mechanism for proactive FT for arbitrary MPI applications. It leverages virtualization techniques combined with health monitoring and load-based migration. We exploit Xen`s live migration mechanism for a guest operating system (OS) to migrate an MPI task from a health-deteriorating node to a healthy one without stopping the MPI task during most of the migration. Our proactive FT daemon orchestrates the tasks of health monitoring, load determination and initiation of guest OS migration. Experimental results demonstrate that live migration hides migration costs and limits the overhead to only a few seconds making it an attractive approach to realize FT in HPC systems. Overall, our enhancements make proactive FT a valuable asset for long-running MPI application that is complementary to reactive FT using full checkpoint\/restart schemes since checkpoint frequencies can be reduced as fewer unanticipated failures are encountered. In the context of OS virtualization, we believe that this is the first comprehensive study of proactive fault tolerance where live migration is actually triggered by health monitoring.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/nagarajan07proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/nagarajan07proactive.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#nagarajan07proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>On Programming Models for Service-Level High Availability<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ares-conference.eu\/ares2007\" target=\"www.ares-conference.eu\/ares2007\">2nd International Conference on Availability, Reliability and Security (ARES) 2007<\/a><\/i>, pages 999-1006, Vienna, Austria, April 10-13, 2007. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2775-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ARES.2007.109\" target=\"publication\">10.1109\/ARES.2007.109<\/a>. Acceptance rate 28.3% (60\/212). <a href=\"javascript:showAbstract('This paper provides an overview of existing programming models for service-level high availability and investigates their differences, similarities, advantages, and disadvantages. Its goal is to help to improve reuse of code and to allow adaptation to quality of service requirements by using a uniform programming model description. It further aims at encouraging a discussion about these programming models and their provided quality of service, such as availability, performance, serviceability, usability, and applicability. Within this context, the presented research focuses on providing high availability for services running on head and service nodes of high-performance computing systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07programming.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07programming.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07programming\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>A Job Pause Service under LAM\/MPI+BLCR for Transparent Fault Tolerance<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ipdps.org\/ipdps2007\" target=\"www.ipdps.org\/ipdps2007\">21st IEEE International Parallel and Distributed Processing Symposium (IPDPS) 2007<\/a><\/i>, pages 1-10, Long Beach, CA, USA, March 26-30, 2007. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-59593-768-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/IPDPS.2007.370307\" target=\"publication\">10.1109\/IPDPS.2007.370307<\/a>. Acceptance rate 26% (109\/419). <a href=\"javascript:showAbstract('Checkpoint\/restart (C\/R) has become a requirement for long-running jobs in large-scale clusters due to a mean-time-to-failure (MTTF) in the order of hours. After a failure, C\/R mechanisms generally require a complete restart of an MPI job from the last checkpoint. A complete restart, however, is unnecessary since all but one node are typically still alive. Furthermore, a restart may result in lengthy job requeuing even though the original job had not exceeded its time quantum. In this paper, we overcome these shortcomings. Instead of job restart, we have developed a transparent mechanism for job pause within LAM\/MPI+BLCR. This mechanism allows live nodes to remain active and roll back to the last checkpoint while failed nodes are dynamically replaced by spares before resuming from the last checkpoint. Our methodology includes LAM\/MPI enhancements in support of scalable group communication with fluctuating number of nodes, reuse of network connections, transparent coordinated checkpoint scheduling and a BLCR enhancement for job pause. Experiments in a cluster with the NAS Parallel Benchmark suite show that our overhead for job pause is comparable to that of a complete job restart. A minimal overhead of 5.6% is only incurred in case migration takes place while the regular checkpoint overhead remains unchanged. Yet, our approach alleviates the need to reboot the LAM run-time environment, which accounts for considerable overhead resulting in net savings of our scheme in the experiments. Our solution further provides full transparency and automation with the additional benefit of reusing existing resources. Executing continues after failures within the scheduled job, \\em \\textiti.e., the application staging overhead is not incurred again in contrast to a restart. Our scheme offers additional potential for savings through incremental checkpointing and proactive diskless live migration, which we are currently working on.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang07job.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/wang07job.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#wang07job\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Kai Uhlemann, Christian Engelmann, and Stephen L. Scott. <b>JOSHUA: Symmetric Active\/Active Replication for Highly Available HPC Job and Resource Management<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster2006.org\" target=\"cluster2006.org\">8th IEEE International Conference on Cluster Computing (Cluster) 2006<\/a><\/i>, pages 1-10, Barcelona, Spain, September 25-28, 2006. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 1-4244-0328-6. ISSN 1552-5244. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTR.2006.311855\" target=\"publication\">10.1109\/CLUSTR.2006.311855<\/a>. Acceptance rate 33.1% (42\/127). <a href=\"javascript:showAbstract('Most of today`s HPC systems employ a single head node for control, which represents a single point of failure as it interrupts an entire HPC system upon failure. Furthermore, it is also a single point of control as it disables an entire HPC system until repair. One of the most important HPC system service running on the head node is the job and resource management. If it goes down, all currently running jobs loose the service they report back to. They have to be restarted once the head node is up and running again. With this paper, we present a generic approach for providing symmetric active\/active replication for highly available HPC job and resource management. The JOSHUA solution provides a virtually synchronous environment for continuous availability without any interruption of service and without any loss of state. Replication is performed externally via the PBS service interface without the need to modify any service code. Test results as well as availability analysis of our proof-of-concept prototype implementation show that continuous availability can be provided by JOSHUA with an acceptable performance trade-off.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/uhlemann06joshua.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/uhlemann06joshua.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#uhlemann06joshua\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Ronald Baumann, Christian Engelmann, and George A. (Al) Geist. <b>A Parallel Plug-in Programming Paradigm<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/hpcc06.lrr.in.tum.de\" target=\"hpcc06.lrr.in.tum.de\">7th International Conference on High Performance Computing and Communications (HPCC) 2006<\/a><\/i>, pages 823-832, Munich, Germany, September 13-15, 2006. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-540-39368-9. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/11847366_85\" target=\"publication\">10.1007\/11847366_85<\/a>. <a href=\"javascript:showAbstract('Software component architectures allow assembly of applications from individual software modules based on clearly defined programming interfaces, thus improving the reuse of existing solutions and simplifying application development. Furthermore, the plug-in programming paradigm additionally enables runtime reconfigurability, making it possible to adapt to changing application needs, such as different application phases, and system properties, like resource availability, by loading\/unloading appropriate software modules. Similar to parallel programs, parallel plug-ins are an abstraction for a set of cooperating individual plug-ins within a parallel application utilizing a software component architecture. Parallel programming paradigms apply to parallel plug-ins in the same way they apply to parallel programs. The research presented in this paper targets the clear definition of parallel plug-ins and the development of a parallel plug-in programming paradigm.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/baumann06parallel.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/baumann06parallel.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#baumann06parallel\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Jyothish Varma, Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Scalable, Fault-Tolerant Membership for MPI Tasks on HPC Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ics-conference.org\/2006\" target=\"www.ics-conference.org\/2006\">20th ACM International Conference on Supercomputing (ICS) 2006<\/a><\/i>, pages 219-228, Cairns, Australia, June 28-30, 2006. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 1-59593-282-8. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1183401.1183433\" target=\"publication\">10.1145\/1183401.1183433<\/a>. Acceptance rate 26.2% (37\/141). <a href=\"javascript:showAbstract('Reliability is increasingly becoming a challenge for high-performance computing (HPC) systems with thousands of nodes, such as IBM`s Blue Gene\/L. A shorter mean-time-to-failure can be addressed by adding fault tolerance to reconfigure working nodes to ensure that communication and computation can progress. However, existing approaches fall short in providing scalability and small reconfiguration overhead within the fault-tolerant layer. This paper contributes a scalable approach to reconfigure the communication infrastructure after node failures. We propose a decentralized (peer-to-peer) protocol that maintains a consistent view of active nodes in the presence of faults. Our protocol shows response times in the order of hundreds of microseconds and single-digit milliseconds for  reconfiguration using MPI over Blue Gene\/L and TCP over  Gigabit, respectively. The protocol can be adapted to match the network topology to further increase performance. We also verify experimental results against a performance model, which demonstrates the scalability of the approach. Hence, the membership service is suitable for deployment in the communication layer of MPI runtime systems, and we have integrated an early version into LAM\/MPI.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/varma06scalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/varma06scalable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#varma06scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Daniel I. Okunbor, Christian Engelmann, and Stephen L. Scott. <b>Exploring Process Groups for Reliability, Availability and Serviceability of Terascale Computing Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.atiner.gr\/docs\/2006AAAPROGRAM_COMP.htm\" target=\"www.atiner.gr\/docs\/2006AAAPROGRAM_COMP.htm\">2nd International Conference on Computer Science and Information Systems 2006<\/a><\/i>, Athens, Greece, June 19-21, 2006. <a href=\"javascript:showAbstract('This paper presents various aspects of reliability, availability and serviceability (RAS) systems as they relate to group communication service, including reliable and total order multicast\/broadcast, virtual synchrony, and failure detection. While the issue of availability, particularly high availability using replication-based architectures has recently received upsurge research interests, much still have to be done in understanding the basic underlying concepts for achieving RAS systems, especially in high-end and high performance computing (HPC) communities. Various attributes of group communication service and the prototype of symmetric active replication following ideas utilized in the Newtop protocol will be discussed. We explore the application of group communication service for RAS HPC, laying the groundwork for its integrated model.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/okunbor06exploring.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#okunbor06exploring\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Kshitij Limaye, Chokchai (Box) Leangsuksun, Zeno Greenwood, Stephen L. Scott, Christian Engelmann, Richard M. Libby, and Kasidit Chanchio. <b>Job-Site Level Fault Tolerance for Cluster and Grid Environments<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster2005.org\" target=\"cluster2005.org\">7th IEEE International Conference on Cluster Computing (Cluster) 2005<\/a><\/i>, pages 1-9, Boston, MA, USA, September 26-30, 2005. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7803-9486-0. ISSN 1552-5244. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTR.2005.347043\" target=\"publication\">10.1109\/CLUSTR.2005.347043<\/a>. Acceptance rate 39.6% (45\/138). <a href=\"javascript:showAbstract('In order to adopt high performance clusters and Grid computing for mission critical applications, fault tolerance is a necessity. Common fault tolerance techniques in distributed systems are normally achieved with checkpoint-recovery and job replication on alternative resources, in cases of a system outage. The first approach depends on the system`s MTTR while the latter approach depends on the availability of alternative sites to run replicas. There is a need for complementing these approaches by proactively handling failures at a job-site level, ensuring the system high availability with no loss of user submitted jobs. This paper discusses a novel fault tolerance technique  that enables the job-site recovery in Beowulf cluster-based grid environments, whereas existing techniques give up a failed system by seeking alternative resources. Our results suggest sizable aggregate performance improvement during an implementation of our method in Globus-enabled HA-OSCAR. The technique called Smart Failover provides a transparent and graceful recovery mechanism that saves job states in a local job-manager queue and transfers those states to the backup server periodically, and in critical system events. Thus whenever a failover occurs, the backup server is able to restart the jobs from their last saved state.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/limaye05jobsite.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#limaye05jobsite\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Hertong Song, Chokchai (Box) Leangsuksun, Raja Nassar, Yudan Liu, Christian Engelmann, and Stephen L. Scott. <b>UML-based Beowulf Cluster Availability Modeling<\/b>. In <i><a href=\"http:\/\/www.world-academy-of-science.org\/IMCSE2005\/ws\/SERP\" target=\"www.world-academy-of-science.org\/IMCSE2005\/ws\/SERP\">International Conference on Software Engineering Research and Practice (SERP) 2005<\/a><\/i>, pages 161-167, Las Vegas, NV, USA, June 27-30, 2005. CSREA Press. ISBN 1-932415-49-1. <a href=\"?page_id=55#song05umlbased\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>Super-Scalable Algorithms for Computing on 100,000 Processors<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.iccs-meeting.org\/iccs2005\" target=\"www.iccs-meeting.org\/iccs2005\">5th International Conference on Computational Science (ICCS) 2005<\/a>, Part I<\/i>, pages 313-320, Atlanta, GA, USA, May 22-25, 2005. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-540-26032-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/11428831_39\" target=\"publication\">10.1007\/11428831_39<\/a>. Acceptance rate 35%. <a href=\"javascript:showAbstract('In the next five years, the number of processors in high-end systems for scientific computing is expected to rise to tens and even hundreds of thousands. For example, the IBM Blue Gene\/L can have up to 128,000 processors and the delivery of the first system is scheduled for 2005. Existing deficiencies in scalability and fault-tolerance of scientific applications need to be addressed soon. If the number of processors grows by a magnitude and efficiency drops by a magnitude, the overall effective computing performance stays the same. Furthermore, the mean time to interrupt of high-end computer systems decreases with scale and complexity. In a 100,000-processor system, failures may occur every couple of minutes and traditional checkpointing may no longer be feasible. With this paper, we summarize our recent research in super-scalable algorithms for computing on 100,000 processors. We introduce the algorithm properties of scale invariance and natural fault tolerance, and discuss how they can be applied to two different classes of algorithms. We also describe a super-scalable diskless checkpointing algorithm for problems that can`t be transformed into a super-scalable variant, or where other solutions are more efficient. Finally, a 100,000-processor simulator is presented as a platform for testing and experimentation.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05superscalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05superscalable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05superscalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"workshops\"><\/a>Peer-reviewed Workshop Papers<\/h4>\n<ol>\n<li>Christian Engelmann, Andrew Ayres, Stephen DeWitt, Michael J. Brim, and Brett Eiffert. <b>Building Resilient Self-Driving Laboratories with the INTERSECT Federated Ecosystem<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc26.supercomputing.org\" target=\"sc26.supercomputing.org\">39th International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2026<\/a>: <a href=\"http:\/\/wordpress.cels.anl.gov\/xloop-2026\/\" target=\"wordpress.cels.anl.gov\/xloop-2026\/\">8th Annual Workshop on Extreme-Scale Experiment-in-the-Loop Computing (XLOOP) 2026<\/a><\/i>, Chicago, IL, USA, November 15, 2026. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. To appear. <a href=\"javascript:showAbstract('Failure resilience in federated ecosystems for instrument science presents a critical challenge. Failures disrupt experiments and make them potentially useless, wasting valuable resources and creating setbacks. Oak Ridge National Laboratory&amp;#39;s Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) offers a federated ecosystem for instrument science, enabling autonomous experiments, self-driving laboratories, smart manufacturing, and AI-driven design, discovery, and evaluation. This paper documents the recent advances in creating a resilient INTERSECT ecosystem. The proposed solution includes a resilient architecture with resilience design patterns, a resilient system of systems (SoS) architecture, and a resilient microservices architecture; and a resilient software development kit with reliable service communication and asynchronous and synchronous failure detection and notification. The resilience capabilities are demonstrated for an autonomous additive manufacturing process with a real-time feedback loop.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#engelmann26building\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen Boehm, Craig A. Bridges, Patrick Widener, Terry Jones, Sheikh Ghafoor, Christian Engelmann, and Olga Kuchar. <b>The INTERSECT Scientific Data Layer: An Ontological Framework for Data Provenance for Complex Scientific Workflows<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/2026.euro-par.org\" target=\"2026.euro-par.org\">32nd European Conference on Parallel and Distributed Computing (Euro-Par) 2026 Workshops<\/a>: <a href=\"http:\/\/www.hipes-workshop.org\/\" target=\"www.hipes-workshop.org\/\">3rd Workshop on High-Performance eScience Tools and Applications (HiPES)<\/a><\/i>, Pisa, Italy, August 25, 2026. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. To appear. <a href=\"javascript:showAbstract('Complex scientific workflows have multiple stages, including experiments, simulations, data analyses, and visualization, generating data and metadata stored across heterogeneous storage infrastructures and used in downstream stages or future experimental campaigns. This paper presents the design and implementation of a comprehensive ontological framework for managing scientific data generated within such workflows, addressing the key challenges of data interoperability, provenance capture, and adherence to Findable, Accessible, Interoperable, and Reusable (FAIR) data principles. The Autonomous Chemistry Laboratory (ACL) at Oak Ridge National Laboratory (ORNL) enables automated liquid phase and solid state synthesis and related chemical analysis. Our framework has been deployed within the ACL for a native and machine-interpretable semantic representation of its ecosystem, including instrument capabilities, synthesis workflows, analytical observations, and experimental results. The proposed approach enables end-to-end provenance tracking, supports heterogeneous data formats, and establishes a foundation for Artificial Intelligence (AI)-ready scientific discovery. We validate the framework through concrete modeling examples.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#boehm28intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Olivera Kotevska, Trong Nguyen, Rafael Ferreira da Silva, Christian Engelmann, and Prasanna Balaprakash. <b>Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/2026.eurosys.org\/\" target=\"2026.eurosys.org\/\">21st European Conference on Computer Systems (EuroSyS)<\/a>: <a href=\"http:\/\/euromlsys.eu\/\" target=\"euromlsys.eu\/\">6th European Workshop on Machine Learning and Systems (EuroMLSys)<\/a><\/i>, pages 439-446, Edinburgh, United Kingdom, April 27, 2026. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 979-8-4007-2605-7. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3805621.3807639\" target=\"publication\">10.1145\/3805621.3807639<\/a>. Acceptance rate 69.2% (18\/26). <a href=\"javascript:showAbstract('Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronization and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kotevska26scalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kotevska26scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Michael J. Brim, Lance Drane, Marshall McDonnell, Christian Engelmann, and Addi Malviya Thakur. <b>A Microservices Architecture Toolkit for Interconnected Science Ecosystems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc24.supercomputing.org\" target=\"sc24.supercomputing.org\">37th International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2024<\/a>: <a href=\"http:\/\/works-workshop.org\/\" target=\"works-workshop.org\/\">19th Workshop on Workflows in Support of Large-Scale  Science (WORKS) 2024<\/a><\/i>, pages 2072-2079, Atlanta, GA, USA, November 18, 2024. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 979-8-3503-5554-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SCW63240.2024.00259\" target=\"publication\">10.1109\/SCW63240.2024.00259<\/a>. Acceptance rate 66.7% (10\/15). <a href=\"javascript:showAbstract('Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for integrating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architecture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability development, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/brim24microservices.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#brim24microservices\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar and Christian Engelmann. <b>RDPM: An Extensible Tool for Resilience Design Patterns Modeling<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/2021.euro-par.org\" target=\"2021.euro-par.org\">27th European Conference on Parallel and Distributed Computing (Euro-Par) 2021 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\">14th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 283-297, Lisbon, Portugal, August 30, 2021. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-031-06155-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-031-06156-1_23\" target=\"publication\">10.1007\/978-3-031-06156-1_23<\/a>. Acceptance rate 66.7% (4\/6). <a href=\"javascript:showAbstract('Resilience to faults, errors, and failures in extreme-scale HPC systems is a critical challenge. Resilience design patterns offer a new, structured hardware and software design approach for improving resilience. While prior work focused on developing performance, reliability, and availability models for resilience design patterns, this paper extends it by providing a Resilience Design Patterns Modeling (RDPM) tool which allows (1) exploring performance, reliability, and availability of each resilience design pattern, (2) offering customization of parameters to optimize performance, reliability, and availability, and (3) allowing investigation of trade-off models for combining multiple patterns for practical resilience solutions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar21rdpm.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar21rdpm\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar and Christian Engelmann. <b>Models for Resilience Design Patterns<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc20.supercomputing.org\" target=\"sc20.supercomputing.org\">33rd International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2020<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2020\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2020\">10th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2020<\/a><\/i>, pages 21-30, Atlanta, GA, USA, November 11, 2020. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7381-1080-6. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS51974.2020.00008\" target=\"publication\">10.1109\/FTXS51974.2020.00008<\/a>. Acceptance rate 66.7% (6\/9). <a href=\"javascript:showAbstract('Resilience plays an important role in supercomputers by providing correct and efficient operation in case of faults, errors, and failures. Resilience design patterns offer blueprints for effectively applying resilience technologies. Prior work focused on developing initial efficiency and performance models for resilience design patterns. This paper extends it by (1) describing performance, reliability, and availability models for all structural resilience design patterns, (2) providing more detailed models that include flowcharts and state diagrams, and (3) introducing the Resilience Design Pattern Modeling (RDPM) tool that calculates and plots the performance, reliability, and availability metrics of individual patterns and pattern combinations.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar20models.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/kumar20models.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#kumar20models\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Piyush Sao, Christian Engelmann, Srinivas Eswar, Oded Green, and Richard Vuduc. <b>Self-stabilizing Connected Components<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc19.supercomputing.org\" target=\"sc19.supercomputing.org\">32nd International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2019<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2019\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2019\">9th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2019<\/a><\/i>, pages 50-59, Denver, CO, USA, November 22, 2019. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-6013-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS49593.2019.00011\" target=\"publication\">10.1109\/FTXS49593.2019.00011<\/a>. Acceptance rate 60.0% (6\/10). <a href=\"javascript:showAbstract('For the problem of computing the connected components of a graph, this paper considers the design of algorithms that are resilient to transient hardware faults, like bit flips. More specifically, it applies the technique of \\emphself-stabilization. A system is self-stabilizing if, when starting from a valid or invalid state, it is guaranteed to reach a valid state after a finite number of steps. Therefore on a machine subject to a transient fault, a self-stabilizing algorithm could recover if that fault caused the system to enter an invalid state. We give a comprehensive analysis of the valid and invalid states during label propagation and derive algorithms to verify and correct the invalid state. The self-stabilizing label-propagation algorithm performs &amp;#36;\\bigoV &amp;#322;og V additional computation and requires \\bigoV additional storage over its conventional counterpart (and, as such, does not increase asymptotic complexity over conventional). When run against a battery of simulated fault injection tests, the self-stabilizing label propagation algorithm exhibits more resilient behavior than a triple modular redundancy (TMR) based fault-tolerant algorithm in 80% of cases. From a performance perspective, it also outperforms TMR as it requires fewer iterations in total. Beyond the fault-tolerance properties of self-stabilizing label-propagation, we believe, they are useful from the theoretical perspective; and may have other use-cases.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/sao19self-stabilizing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/sao19self-stabilizing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#sao19self-stabilizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Geoffroy R. Vall&eacute;e, and Swaroop Pophale. <b>Concepts for OpenMP Target Offload Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/parallel.auckland.ac.nz\/iwomp2019\" target=\"parallel.auckland.ac.nz\/iwomp2019\">15th International Workshop on OpenMP (IWOMP) 2019<\/a><\/i>, pages 78-93, Auckland, New Zealand, September 11-13, 2019. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-030-28595-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-28596-8_6\" target=\"publication\">10.1007\/978-3-030-28596-8_6<\/a>. <a href=\"javascript:showAbstract('Recent reliability issues with one of the fastest supercomputers in the world, Titan at Oak Ridge National Laboratory, demonstrated the need for resilience in large-scale heterogeneous computing. OpenMP currently does not address error and failure behavior. This paper takes a first step toward resilience for heterogeneous systems by providing the concepts for resilient OpenMP offload to devices. Using real-world error and failure observations, the paper describes the concepts and terminology for resilient OpenMP target offload, including error and failure classes and resilience strategies. It details the experienced general-purpose computing on graphics processing units errors and failures in Titan. It further proposes improvements in OpenMP, including a preliminary prototype design, to support resilient offload to devices for efficient handling of errors and failures in heterogeneous high-performance computing systems');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19concepts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann19concepts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19concepts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>A Comprehensive Informative Metric for Analyzing HPC System Status using the LogSCAN Platform<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc18.supercomputing.org\" target=\"sc18.supercomputing.org\">31st International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\">8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2018<\/a><\/i>, pages 29-38, Dallas, TX, USA, November 16, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-0222-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS.2018.00007\" target=\"publication\">10.1109\/FTXS.2018.00007<\/a>. Acceptance rate 45.0% (9\/20). <a href=\"javascript:showAbstract('Log processing by Spark and Cassandra-based ANalytics (LogSCAN) is a newly developed analytical platform that provides flexible and scalable data gathering, transformation and computation. One major challenge is to effectively summarize the status of a complex computer system, such as the Titan supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). Although there is plenty of operational and maintenance information collected and stored in real time, which may yield insights about short- and long-term system status, it is difficult to present this information in a comprehensive form. In this work, we present system information entropy (SIE), a newly developed metric that leverages the powers of traditional machine learning techniques and information theory. By compressing the multi-variant multi-dimensional event information recorded during the operation of the targeted system into a single time series of SIE, we demonstrate that the historical system status can be sensitively represented concisely and comprehensively. Given a sharp indicator as SIE, we argue that follow-up analytics based on SIE will reveal in-depth knowledge about system status using other sophisticated approaches, such as pattern recognition in the temporal domain or causality analysis incorporating extra independent metrics of the system.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18comprehensive2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hui18comprehensive2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hui18comprehensive2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf and Christian Engelmann. <b>Analyzing the Impact of System Reliability Events on Applications in the Titan Supercomputer<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc18.supercomputing.org\" target=\"sc18.supercomputing.org\">31st International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\">8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2018<\/a><\/i>, pages 39-48, Dallas, TX, USA, November 16, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-0222-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS.2018.00008\" target=\"publication\">10.1109\/FTXS.2018.00008<\/a>. Acceptance rate 45.0% (9\/20). <a href=\"javascript:showAbstract('Extreme-scale computing systems employ Reliability, Availability and Serviceability (RAS) mechanisms and infrastructure to log events from multiple system components. In this paper, we analyze RAS logs in conjunction with the  application placement and scheduling database, in order to  understand the impact of common RAS events on application performance. This study conducted on the records of about 2 million applications executed on Titan supercomputer provides important insights for system users, operators and computer science researchers. In this paper, we investigate the impact of RAS events on application performance and its variability by comparing cases where events are recorded with corresponding cases where no events are recorded. Such a statistical investigation is possible since we observed that system users tend to execute their applications multiple times. Our analysis reveals that most RAS events do impact application performance, although not always. We also find that different system components affect application performance differently. In particular, our investigation includes the following components: parallel file system, processor, memory, graphics processing units, system and user software issues. Our work establishes the importance of providing feedback to system users for increasing operational efficiency of extreme-scale systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18analyzing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18analyzing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18analyzing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Byung Hoon (Hoony) Park, Yawei Hui, Swen Boehm, Rizwan Ashraf, Christian Engelmann, and Christopher Layton. <b>A Big Data Analytics Framework for HPC Log Data: Three Case Studies Using the Titan Supercomputer Log<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster2018.github.io\" target=\"cluster2018.github.io\">19th IEEE International Conference on Cluster Computing (Cluster) 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/hpcmaspa2018\" target=\"sites.google.com\/site\/hpcmaspa2018\">5th Workshop on Monitoring and Analysis for High Performance Systems Plus Applications (HPCMASPA) 2018<\/a><\/i>, pages 571-579, Belfast, UK, September 10, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-8319-4. ISSN 2168-9253. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTER.2018.00073\" target=\"publication\">10.1109\/CLUSTER.2018.00073<\/a>. <a href=\"javascript:showAbstract('Reliability, availability and serviceability (RAS) logs of high performance computing (HPC) resources, when closely investigated in spatial and temporal dimensions, can provide invaluable information regarding system status, performance, and resource utilization. These data are often generated from multiple logging systems and sensors that cover many components of the system. The analysis of these data for finding persistent temporal and spatial insights faces two main difficulties: the volume of RAS logs makes manual inspection difficult and the unstructured nature and unique properties of log data produced by each subsystem adds another dimension of difficulty in identifying implicit correlation among recorded events. To address these issues, we recently developed a multi-user Big Data analytics framework for HPC log data at Oak Ridge National Laboratory (ORNL). This paper introduces three in-progress data analytics projects that leverage this framework to assess system status, mine event patterns, and study correlations between user applications and system events. We describe the motivation of each project and detail their workflows using three years of log data collected from ORNL&amp;#39;s Titan supercomputer.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/park18big.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/park18big.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#park18big\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf and Christian Engelmann. <b>Performance Efficient Multiresilience using Checkpoint Recovery in Iterative Algorithms<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2018.org\" target=\"europar2018.org\">24th European Conference on Parallel and Distributed Computing (Euro-Par) 2018 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2018\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2018\">11th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 813-825, Turin, Italy, August 28, 2018. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-030-10549-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-10549-5_63\" target=\"publication\">10.1007\/978-3-030-10549-5_63<\/a>. Acceptance rate 50.0% (4\/8). <a href=\"javascript:showAbstract('In this paper, we address the design challenge of building multiresilient iterative high-performance computing (HPC) applications. Multiresilience in HPC applications is the ability to tolerate and maintain forward progress in the presence of both soft errors and process failures. We address the challenge by proposing performance models which are useful to design performance efficient and resilient iterative applications. The models consider the interaction between soft error and process failure resilience solutions. We experimented with a linear solver application with two distinct kinds of soft error detectors: one detector is high overhead and high accuracy, whereas the second is low overhead and low accuracy. We show how both can be leveraged for verifying the integrity of checkpointed state used to recover from both soft errors and process failures. Our results show the performance efficiency and resiliency benefit of employing the low overhead detector with high frequency within the checkpoint interval, so that timely soft error recovery can take place, resulting in less re-computed work.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18performance.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18performance.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18performance\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Byung Hoon (Hoony) Park, Saurabh Hukerikar, Christian Engelmann, and Ryan Adamson. <b>Big Data Meets HPC Log Analytics: Scalable Approach to Understanding Systems at Extreme Scale<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster17.github.io\" target=\"cluster17.github.io\">18th IEEE International Conference on Cluster Computing (Cluster) 2017<\/a>: <a href=\"http:\/\/sites.google.com\/site\/hpcmaspa2017\" target=\"sites.google.com\/site\/hpcmaspa2017\">4th Workshop on Monitoring and Analysis for High Performance Systems Plus Applications (HPCMASPA) 2017<\/a><\/i>, pages 758-765, Honolulu, HI, USA, September 5, 2017. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-2327-5. ISSN 2168-9253. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTER.2017.113\" target=\"publication\">10.1109\/CLUSTER.2017.113<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s high-performance computing (HPC) systems are heavily instrumented generating logs containing information about abnormal events, such as critical conditions, faults, errors and failures, system resource utilization, and about the resource usage of user applications. These logs, once fully analyzed and correlated, can produce detailed information about the system health, root causes of failures, and analyze an application's interactions with the system, providing invaluable insights to domain scientists and system administrators. However, processing HPC logs requires deep understanding of hardware and software components at multiple layers of the system stack. Moreover, most log data is unstructured and voluminous, making it more difficult for scientists and engineers to analyze the data. With rapid increases in the scale and complexity of HPC systems, log data processing is becoming a big data challenge. This paper introduces a HPC log data analytics framework that is based on a distributed NoSQL database technology, which provides scalability and high availability, and the Apache Spark for rapid in-memory processing of log data. The framework enables the extraction of a range of information about the system so that system administrators and end users alike can obtain necessary insights for their specific needs. We describe our experience with using this framework to glean insights from the log data derived from the Titan supercomputer at the Oak Ridge National Laboratory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/park17big.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/park17big.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#park17big\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Pattern-based Modeling of High-Performance Computing Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2017.usc.es\" target=\"europar2017.usc.es\">23rd European Conference on Parallel and Distributed Computing (Euro-Par) 2017 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2017\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2017\">10th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 557-568, Santiago de Compostela, Spain, August 29, 2017. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-75177-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-75178-8_45\" target=\"publication\">10.1007\/978-3-319-75178-8_45<\/a>. Acceptance rate 66.7% (4\/6). <a href=\"javascript:showAbstract('The design of supercomputing systems and their applications must consider resilience, and power consumption as the key design parameters when designing to achieve higher performance. In previous work, we established a structured methodology for developing resilience solutions based on the concept of design patterns. In this paper we discuss analytical models for the design patterns to support quantitative analysis of their performance and reliability characteristics.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17pattern-based.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar17pattern-based.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar17pattern-based\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar, Rizwan Ashraf, and Christian Engelmann. <b>Towards New Metrics for High-Performance Computing Resilience<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.hpdc.org\/2017\" target=\"www.hpdc.org\/2017\">26th ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC) 2017<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2017\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2017\">7th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2017<\/a><\/i>, pages 23-30, Washington, D.C., June 26-30, 2017. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-5001-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3086157.3086163\" target=\"publication\">10.1145\/3086157.3086163<\/a>. Acceptance rate 83.3% (5\/6). <a href=\"javascript:showAbstract('Ensuring the reliability of applications is becoming an increasingly important challenge as high-performance computing (HPC) systems experience an ever-growing number of faults, errors and failures. While the HPC community has made substantial progress in developing various resilience solutions, it continues to rely on platform-based metrics to quantify application resiliency improvements. The resilience of an HPC application is concerned with the reliability of the application outcome as well as the fault handling efficiency. To understand the scope of impact, effective coverage and performance efficiency of existing and emerging resilience solutions, there is a need for new metrics. In this paper, we develop new ways to quantify resilience that consider both the reliability and the performance characteristics of the solutions from the perspective of HPC applications. As HPC systems continue to evolve in terms of scale and complexity, it is expected that applications will experience various types of faults, errors and failures, which will require applications to apply multiple resilience solutions across the system stack. The proposed metrics are intended to be useful for understanding the combined impact of these solutions on an application&amp;#39;s ability to produce correct results and to evaluate their overall impact on an application's performance in the presence of various modes of faults.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17towards.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar17towards.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar17towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Language Support for Reliable Memory Regions<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/lcpc2016.wordpress.com\" target=\"lcpc2016.wordpress.com\">29th International Workshop on Languages and Compilers for Parallel Computing<\/a><\/i>, pages 73-87, Rochester, NY, USA, September 28-30, 2016. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-52708-6. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-52709-3_6\" target=\"publication\">10.1007\/978-3-319-52709-3_6<\/a>. Acceptance rate 76.9% (20\/26). <a href=\"javascript:showAbstract('The path to exascale computational capabilities in high-performance computing (HPC) systems is challenged by the evolution of the architectures of supercomputing systems. The constraints of power have driven designs that include increasingly heterogeneous architectures and complex memory hierarchies. These systems are also expected to experience in an increased rate of errors, such that the applications will no longer be able to assume correct behavior of the underlying machine. To enable the scientific community to succeed in scaling their applications and harness the capabilities of exascale systems, we need software strategies that provide mechanisms for explicit management of locality and resilience to errors in the system. In prior work, we introduced the concept of explicitly reliable memory regions, called havens. Memory management using havens supports selective reliability through a region-based approach to memory allocation. Havens enable the creation of explicit software-enabled robust memory containers for which resilient behavior is guaranteed. In this paper, we propose language support for havens through type annotations that make the structure of a program&amp;#39;s havens more explicit. We describe how the extended haven-based memory management model is implemented and the impact on the resiliency of a conjugate gradient application.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar16language.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar16language.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar16language\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Christian Engelmann, Geoffroy Vall&eacute;e, Ferrol Aderholdt, and Stephen L. Scott. <b>A Cooperative Approach to Virtual Machine Based Fault Injection<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2016.inria.fr\" target=\"europar2016.inria.fr\">22nd European Conference on Parallel and Distributed Computing (Euro-Par) 2016 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2016\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2016\">9th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 671-682, Grenoble, France, August 23, 2016. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-58943-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-58943-5_54\" target=\"publication\">10.1007\/978-3-319-58943-5_54<\/a>. Acceptance rate 55.6% (5\/9). <a href=\"javascript:showAbstract('Resilience investigations often employ fault injection (FI) tools to study the effects of simulated errors on a target system. It is important to keep the target system under test (SUT) isolated from the controlling environment in order to maintain control of the experiment. Virtual machines (VMs) have been used to aid these investigations due to the strong isolation properties of system-level virtualization. A key challenge in fault injection tools is to gain proper insight and context about the SUT. In VM-based FI tools, this challenge of target con- text is increased due to the separation between host and guest (VM). We discuss an approach to VM-based FI that leverages virtual machine introspection (VMI) methods to gain insight into the target&amp;#39;s context running within the VM. The key to this environment is the ability to provide basic information to the FI system that can be used to create a map of the target environment. We describe a proof- of-concept implementation and a demonstration of its use to introduce simulated soft errors into an iterative solver benchmark running in user-space of a guest VM.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton16cooperative.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton16cooperative.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton16cooperative\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Zachary Parchman, Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, and David E. Bernholdt. <b>Adding Fault Tolerance to NPB Benchmarks Using ULFM<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.hpdc.org\/2016\" target=\"www.hpdc.org\/2016\">25th ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC) 2016<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2016\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2016\">6th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2016<\/a><\/i>, pages 19-26, Kyoto, Japan, May 31 &#8211; June 4, 2016. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-4349-7. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/2909428.2909429\" target=\"publication\">10.1145\/2909428.2909429<\/a>. Acceptance rate 85.7% (6\/7). <a href=\"javascript:showAbstract('In the world of high-performance computing, fault tolerance and application resilience are becoming some of the primary concerns because of increasing hardware failures and memory corruptions. While the research community has been investigating various options, from system-level solutions to application-level solutions, standards such as the Message Passing Interface (MPI) are also starting at including such capabilities. The current proposal for MPI fault tolerant is centered around the User-Level Failure Mitigation (ULFM) concept, which provides means for fault detection and recovery of the MPI layer. This approach does not address application-level recovery, which is current left to application developers. In this work, we present a modification of some of the benchmarks of the NAS parallel benchmark (NPB) to include support of the ULFM capabilities as well as application- level strategies and mechanisms for application-level failure recovery. As such, we present: (i) an application-level library to &amp;#34;checkpoint&amp;#34; data, (ii) extensions of NPB benchmarks for fault tolerance based on different strategies, (iii) a fault injection tool, and (iv) some preliminary experiments that shows the impact of such fault tolerant strategies on the application execution.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/parchman16adding.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/parchman16adding.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#parchman16adding\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Garry Smith, Christian Engelmann, Geoffroy Vall&eacute;e, Ferrol Aderholdt, and Stephen L. Scott. <b>What is the right balance for performance and isolation with virtualization in HPC?<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2014.dcc.fc.up.pt\" target=\"europar2014.dcc.fc.up.pt\">20th European Conference on Parallel and Distributed Computing (Euro-Par) 2014 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2014\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2014\">7th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 570-581, Porto, Portugal, August 25, 2014. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-14325-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-14325-5_49\" target=\"publication\">10.1007\/978-3-319-14325-5_49<\/a>. Acceptance rate 60.0% (6\/10). <a href=\"javascript:showAbstract('The use of virtualization in high-performance computing (HPC) has been suggested as a means to provide tailored services and added functionality that many users expect from full-featured Linux cluster environments. While the use of virtual machines in HPC can offer several benefits, maintaining performance is a crucial factor. In some instances performance criteria are placed above isolation properties and selective relaxation of isolation for performance is an important characteristic when considering resilience for HPC environments employing virtualization. In this paper we consider some of the factors associated with balancing performance and isolation in configurations that employ virtual machines. In this context, we propose a classification of errors based on the concept of &amp;#34;error zones&amp;#34;, as well as a detailed analysis of the trade-offs between resilience and performance based on the level of isolation provided by virtualization solutions. Finally, the results from a set of experiments are presented, that use different virtualization solutions, and in doing so allow further elucidation of the topic.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton14what.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton14what.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton14what\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>Toward a Performance\/Resilience Tool for Hardware\/Software Co-Design of High-Performance Computing Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icpp2013.ens-lyon.fr\" target=\"icpp2013.ens-lyon.fr\">42nd International Conference on Parallel Processing (ICPP) 2013<\/a>: <a href=\"http:\/\/www.psti-workshop.org\" target=\"www.psti-workshop.org\">4th International Workshop on Parallel Software Tools and Tool Infrastructures (PSTI)<\/a><\/i>, pages 962-971, Lyon, France, October 2, 2013. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-5117-3. ISSN 0190-3918. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICPP.2013.114\" target=\"publication\">10.1109\/ICPP.2013.114<\/a>. <a href=\"javascript:showAbstract('xSim is a simulation-based performance investigation toolkit that permits running high-performance computing (HPC) applications in a controlled environment with millions of concurrent execution threads, while observing application performance in a simulated extreme-scale system for hardware\/software co-design. The presented work details newly developed features for xSim that permit the injection of MPI process failures, the propagation\/detection\/notification of such failures within the simulation, and their handling using application-level checkpoint\/restart. These new capabilities enable the observation of application behavior and performance under failure within a simulated future-generation HPC system using the most common fault handling technique.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13toward.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann13toward.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann13toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mahesh Lagadapati, Frank Mueller, and Christian Engelmann. <b>Tools for Simulation and Benchmark Generation at Exascale<\/b>. In <i>Proceedings of the <a href=\"http:\/\/tools.zih.tu-dresden.de\/2013\/\" target=\"tools.zih.tu-dresden.de\/2013\/\">7th Parallel Tools Workshop<\/a><\/i>, pages 19-24, Dresden, Germany, September 3-4, 2013. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-08143-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-08144-1_2\" target=\"publication\">10.1007\/978-3-319-08144-1_2<\/a>. <a href=\"javascript:showAbstract('The path to exascale high-performance computing (HPC) poses several challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Investigating the performance of parallel applications at scale on future architectures and the performance impact of different architecture choices is an important component of HPC hardware\/software co-design. Simulations using models of future HPC systems and communication traces from applications running on existing HPC systems can offer an insight into the performance of future architectures. This work targets technology developed for scalable application tracing of communication events and memory profiles, but can be extended to other areas, such as I\/O, control flow, and data flow. It further focuses on extreme-scale simulation of millions of Message Passing Interface (MPI) ranks using a lightweight parallel discrete event simulation (PDES) toolkit for performance evaluation. Instead of simply replaying a trace within a simulation, the approach is to generate a benchmark from it and to run this benchmark within a simulation using models to reflect the performance characteristics of future-generation HPC systems. This provides a number of benefits, such as eliminating the data intensive trace replay and enabling simulations at different scales. The presented work utilizes the ScalaTrace tool to generate scalable trace files, the ScalaBenchGen tool to generate the benchmark, and the xSim tool to run the benchmark within a simulation.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/lagadapati13tools.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/lagadapati13tools.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#lagadapati13tools\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Swen B&ouml;hm, Christian Engelmann, and Geoffroy Vall&eacute;e. <b>Using Performance Tools to Support Experiments in HPC Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.europar2013.org\/\" target=\"www.europar2013.org\/\">19th European Conference on Parallel and Distributed Computing (Euro-Par) 2013 Workshops<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2013\" target=\"xcr.cenit.latech.edu\/resilience2013\">6th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 727-736, Aachen, Germany, August 26, 2013. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-54419-4. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-54420-0_71\" target=\"publication\">10.1007\/978-3-642-54420-0_71<\/a>. Acceptance rate 87.5% (7\/8). <a href=\"javascript:showAbstract('The high performance computing (HPC) community is working to address fault tolerance and resilience concerns for current and future large scale computing platforms. This is driving enhancements in the programming environments, specifically research on enhancing message passing libraries to support fault tolerant computing capabilities. The community has also recognized that tools for resilience experimentation are greatly lacking. However, we argue that there are several parallels between &amp;#34;performance tools&amp;#34; and ``resilience tools&amp;#39;'. As such, we believe the rich set of HPC performance-focused tools can be extended (repurposed) to benefit the resilience community. In this paper, we describe the initial motivation to leverage standard HPC performance analysis techniques to aid in developing diagnostic tools to assist fault tolerance experiments for HPC applications. These diagnosis procedures help to provide context for the system when the errors (failures) occurred. We describe our initial work in leveraging an MPI performance trace tool to assist in providing global context during fault injection experiments. Such tools will assist the HPC resilience community as they extend existing and new application codes to support fault tolerances.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton13using.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton13using.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton13using\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Ian S. Jones and Christian Engelmann. <b>Simulation of Large-Scale HPC Architectures<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icpp2011.org\" target=\"icpp2011.org\">40th International Conference on Parallel Processing (ICPP) 2011<\/a>: <a href=\"http:\/\/www.psti-workshop.org\" target=\"www.psti-workshop.org\">2nd International Workshop on Parallel Software Tools and Tool Infrastructures (PSTI)<\/a><\/i>, pages 447-456, Taipei, Taiwan, September 13-19, 2011. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4511-0. ISSN 1530-2016. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICPPW.2011.44\" target=\"publication\">10.1109\/ICPPW.2011.44<\/a>. <a href=\"javascript:showAbstract('The Extreme-scale Simulator (xSim) is a recently developed performance investigation toolkit that permits running high-performance computing (HPC) applications in a controlled environment with millions of concurrent execution threads. It allows observing parallel application performance properties in a simulated extreme-scale HPC system to further assist in HPC hardware and application software co-design on the road toward multi-petascale and exascale computing. This paper presents a newly implemented network model for the xSim performance investigation toolkit that is capable of providing simulation support for a variety of HPC network architectures with the appropriate trade-off between simulation scalability and accuracy. The taken approach focuses on a scalable distributed solution with latency and bandwidth restrictions for the simulated network. Different network architectures, such as star, ring, mesh, torus, twisted torus and tree, as well as hierarchical combinations, such as to simulate network-on-chip and network-on-node, are supported. Network traffic congestion modeling is omitted to gain simulation scalability by reducing simulation accuracy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/jones11simulation.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/jones11simulation.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#jones11simulation\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Kurt Ferreira, Frank Mueller, and Christian Engelmann. <b>A Tunable, Software-based DRAM Error Detection and Correction Library for HPC<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2011.bordeaux.inria.fr\/\" target=\"europar2011.bordeaux.inria.fr\/\">17th European Conference on Parallel and Distributed Computing (Euro-Par) 2011 Workshops, Part II<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2011\" target=\"xcr.cenit.latech.edu\/resilience2011\">4th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 251-261, Bordeaux, France, August 29 &#8211; September 2, 2011. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-29740-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-29740-3_29\" target=\"publication\">10.1007\/978-3-642-29740-3_29<\/a>. Acceptance rate 60.0% (12\/20). <a href=\"javascript:showAbstract('Proposed exascale systems will present a number of considerable resiliency challenges. In particular, DRAM soft-errors, or bit-flips, are expected to greatly increase due to the increased memory density of these systems. Current hardware-based fault-tolerance methods will be unsuitable for addressing the expected soft error frequency rate. As a result, additional software will be needed to address this challenge. In this paper we introduce LIBSDC, a tunable, transparent silent data corruption detection and correction library for HPC applications. LIBSDC provides comprehensive SDC protection for program memory by implementing on-demand page integrity verification. Experimental benchmarks with Mantevo HPCCG show that once tuned, LIBSDC is able to achieve SDC protection with 50% overhead of resources, less than the 100% needed for double modular redundancy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala11tunable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#fiala11tunable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Geoffroy R. Vall&eacute;e, Christian Engelmann, and Stephen L. Scott. <b>A Case for Virtual Machine based Fault Injection in a High-Performance Computing Environment<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2011.bordeaux.inria.fr\/\" target=\"europar2011.bordeaux.inria.fr\/\">17th European Conference on Parallel and Distributed Computing (Euro-Par) 2011<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/hpcvirt2011\" target=\"www.csm.ornl.gov\/srt\/conferences\/hpcvirt2011\">5th Workshop on System-level Virtualization for High Performance Computing (HPCVirt)<\/a><\/i>, pages 234-243, Bordeaux, France, August 29 &#8211; September 2, 2011. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-29737. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-29737-3_27\" target=\"publication\">10.1007\/978-3-642-29737-3_27<\/a>. <a href=\"javascript:showAbstract('Large-scale computing platforms provide tremendous capabilities for scientific discovery. These systems have hundreds of thousands of computing cores, hundreds of terabytes of memory, and enormous high-performance interconnection networks. These systems are facing enormous challenges to achieve performance at such scale. Failures are an Achilles heel of these enormous systems. As applications and system software scale up to multi-petaflop and beyond to exascale platforms, the occurrence of failure will be much more common. This has given rise to a push in fault-tolerance and resilience research for HPC systems. This includes work on log analysis to identify types of failures, enhancements to the Message Passing Interface (MPI) to incorporate fault awareness, and a variety of fault tolerance mechanisms that span redundant computation, algorithm based fault tolerance, and advanced checkpoint\/ restart techniques. While there is much work to be done on the FT\/Resilience mechanisms for such large-scale systems, there is also a profound gap in the tools for experimentation. This gap is compounded by the fact that HPC environments have stringent performance requirements and are often highly customized. The tool chain for these systems are often tailored for the platform and while the majority of systems on the Top500 Supercomputer list run Linux, these operating environments typically contain many site\/machine specific enhancements. Therefore, it is desirable to maintain a consistent execution environment to minimize end-user (scientist) interruption. The work on system-level virtualization for HPC system offers a unique opportunity to maintain a consistent execution environment via a virtual machine (VM). Recent work on virtualization for HPC has shown that low-overhead, high performance systems can be realized [1, 2] Virtualization also provides a clean abstraction for building experimental tools for investigation into the effects of failures in HPC and the related research on FT\/ Resilience mechanisms and policies. In this paper we discuss the motivation for tools to perform fault injection in an HPC context, and outline an approach that can leverage virtualization.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton11case.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton11case.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton11case\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Frank Lauer. <b>Facilitating Co-Design for Extreme-Scale Systems Through Lightweight Simulation<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.cluster2010.org\" target=\"www.cluster2010.org\">12th IEEE International Conference on Cluster Computing (Cluster) 2010<\/a>: <a href=\"http:\/\/www2.wmin.ac.uk\/getovv\/aacec10.html\" target=\"www2.wmin.ac.uk\/getovv\/aacec10.html\">1st Workshop on Application\/Architecture Co-design for Extreme-scale Computing (AACEC)<\/a><\/i>, pages 1-8, Hersonissos, Crete, Greece, September 20-24, 2010. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4244-8395-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTERWKSP.2010.5613113\" target=\"publication\">10.1109\/CLUSTERWKSP.2010.5613113<\/a>. <a href=\"javascript:showAbstract('This work focuses on tools for investigating algorithm performance at extreme scale with millions of concurrent threads and for evaluating the impact of future architecture choices to facilitate the co-design of high-performance computing (HPC) architectures and applications. The approach focuses on lightweight simulation of extreme-scale HPC systems with the needed amount of accuracy. The prototype presented in this paper is able to provide this capability using a parallel discrete event simulation (PDES), such that a Message Passing Interface (MPI) application can be executed at extreme scale, and its performance properties can be evaluated. The results of an initial prototype are encouraging as a simple hello world MPI program could be scaled up to 1,048,576 virtual MPI processes on a four-node cluster, and the performance properties of two MPI programs could be evaluated at up to 1,024 and 16,384 virtual MPI processes on the same system.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann10facilitating.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann10facilitating.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10facilitating\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>George Ostrouchov, Thomas Naughton, Christian Engelmann, Geoffroy R. Vall&eacute;e, and Stephen L. Scott. <b>Nonparametric Multivariate Anomaly Analysis in Support of HPC Resilience<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.oerc.ox.ac.uk\/ieee\" target=\"www.oerc.ox.ac.uk\/ieee\">5th IEEE International Conference on e-Science (e-Science) 2009<\/a>: <a href=\"http:\/\/www.oerc.ox.ac.uk\/ieee\/workshops\/workshops\/computational-science\" target=\"www.oerc.ox.ac.uk\/ieee\/workshops\/workshops\/computational-science\">Workshop on Computational Science<\/a><\/i>, pages 80-85, Oxford, UK, December 9-11, 2009. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4244-5946-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ESCIW.2009.5407992\" target=\"publication\">10.1109\/ESCIW.2009.5407992<\/a>. <a href=\"javascript:showAbstract('Large-scale computing systems provide great potential for scientific exploration. However, the complexity that accompanies these enormous machines raises challeges for both, users and operators. The effective use of such systems is often hampered by failures encountered when running applications on systems containing tens-of-thousands of nodes and hundreds-of-thousands of compute cores capable of yielding petaflops of performance. In systems of this size failure detection is complicated and root-cause diagnosis difficult. This paper describes our recent work in the identification of anomalies in monitoring data and system logs to provide further insights into machine status, runtime behavior, failure modes and failure root causes. It discusses the details of an initial prototype that gathers the data and uses statistical techniques for analysis.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ostrouchov09nonparametric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ostrouchov09nonparametric.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ostrouchov09nonparametric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Wesley Bland, Geoffroy R. Vall&eacute;e, Christian Engelmann, and Stephen L. Scott. <b>Fault Injection Framework for System Resilience Evaluation &#8211; Fake Faults for Finding Future Failures<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.lrz-muenchen.de\/hpdc2009\" target=\"www.lrz-muenchen.de\/hpdc2009\">18th International Symposium on High Performance Distributed Computing (HPDC) 2009<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2009\" target=\"xcr.cenit.latech.edu\/resilience2009\">2nd Workshop on Resiliency in High Performance Computing (Resilience) 2009<\/a><\/i>, pages 23-28, Munich, Germany, June 9, 2009. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-60558-587-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1552526.1552530\" target=\"publication\">10.1145\/1552526.1552530<\/a>. <a href=\"javascript:showAbstract('As high-performance computing (HPC) systems increase in size and complexity they become more difficult to manage. The enormous component counts associated with these large systems lead to significant challenges in system reliability and availability. This in turn is driving research into the resilience of large scale systems, which seeks to curb the effects of increased failures at large scales by masking the inevitable faults in these systems. The basic premise being that failure must be accepted as a reality of large scale system and coped with accordingly through system resilience. A key component in the development and evaluation of system resilience techniques is having a means to conduct controlled experiments. A common method for performing such experiments is to generate synthetic faults and study the resulting effects. In this paper we discuss the motivation and our initial use of software fault injection to support the evaluation of resilience for HPC systems. We mention background and related work in the area and discuss the design of a tool to aid in fault injection experiments for both user-space (application-level) and system-level failures.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton09fault.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton09fault.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton09fault\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Anand Tikotekar, Hong H. Ong, Sadaf Alam, Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, and Stephen L. Scott. <b>Performance Comparison of Two Virtual Machine Scenarios Using an HPC Application &#8211; A Case study Using Molecular Dynamics Simulations<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.csm.ornl.gov\/srt\/hpcvirt09\" target=\"www.csm.ornl.gov\/srt\/hpcvirt09\">3rd Workshop on System-level Virtualization for High Performance Computing (HPCVirt) 2009<\/a>, in conjunction with the <a href=\"http:\/\/www.eurosys.org\/2009\" target=\"www.eurosys.org\/2009\">4th ACM SIGOPS European Conference on Computer Systems (EuroSys) 2009<\/a><\/i>, pages 33-40, Nuremberg, Germany, March 30, 2009. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-60558-465-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1519138.1519143\" target=\"publication\">10.1145\/1519138.1519143<\/a>. <a href=\"javascript:showAbstract('Obtaining high flexibility to performance-loss ratio is a key challenge of today&amp;#39;s HPC virtual environment landscape. And while extensive research has been targeted at extracting more performance from virtual machines, the idea that whether novel virtual machine usage scenarios could lead to high flexibility Vs performance trade-off has received less attention. We, in this paper, take a step forward by studying and comparing the performance implications of running the Large-scale Atomic\/Molecular Massively Parallel Simulator (LAMMPS) application on two virtual machine configurations. First configuration consists of two virtual machines per node with 1 application process per virtual machine. The second configuration consists of 1 virtual machine per node with 2 processes per virtual machine. Xen has been used as an hypervisor and standard Linux as a guest virtual machine. Our results show that the difference in overall performance impact on LAMMPS between the two virtual machine configurations described above is around 3%. We also study the difference in performance impact in terms of each configuration's individual metrics such as CPU, I\/O, Memory, and interrupt\/context switches.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tikotekar09performance.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tikotekar09performance.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tikotekar09performance\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy R. Vall&eacute;e, Thomas Naughton, Hong H. Ong, Anand Tikotekar, Christian Engelmann, Wesley Bland, Ferrol Aderholt, and Stephen L. Scott. <b>Virtual System Environments<\/b>. In <i>Communications in Computer and Information Science: Proceedings of the <a href=\"http:\/\/www.dmtf.org\/svm08\" target=\"www.dmtf.org\/svm08\">2nd DMTF Academic Alliance Workshop on Systems and Virtualization Management: Standards and New Technologies (SVM) 2008<\/a><\/i>, pages 72-83, Munich, Germany, October 21-22, 2008. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-540-88707-2. ISSN 1865-0929. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-540-88708-9_7\" target=\"publication\">10.1007\/978-3-540-88708-9_7<\/a>. <a href=\"javascript:showAbstract('Distributed and parallel systems are typically managed with static settings: the operating system (OS) and the runtime environment (RTE) are specified at a given time and cannot be changed to fit an application`s needs. This means that every time application developers want to use their application on a new execution platform, the application has to be ported to this new environment, which may be expensive in terms of application modifications and developer time. However, the science resides in the applications and not in the OS or the RTE. Therefore, it should be beneficial to adapt the OS and the RTE to the application instead of adapting the applications to the OS and the RTE. This document presents the concept of Virtual System Environments (VSE), which enables application developers to specify and create a virtual environment that properly fits their application`s needs. For that four challenges have to be addressed: (i) definition of the VSE itself by the application developers, (ii) deployment of the VSE, (iii) system administration for the platform, and (iv) protection of the platform from the running VSE. We therefore present an integrated tool for the definition and deployment of VSEs on top of traditional and virtual (i.e., using system-level virtualization) execution platforms. This tool provides the capability to choose the degree of delegation for system administration tasks and the degree of protection from the application (e.g., using virtual machines). To summarize, the VSE concept enables the customization of the OS\/RTE used for the execution of application by users without compromising local system administration rules and execution platform protection constraints.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/vallee08virtual.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#vallee08virtual\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Anand Tikotekar, Geoffroy Vall&eacute;e, Thomas Naughton, Hong H. Ong, Christian Engelmann, and Stephen L. Scott. <b>An Analysis of HPC Benchmark Applications in Virtual Machine Environments<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2008.caos.uab.es\" target=\"europar2008.caos.uab.es\">14th European Conference on Parallel and Distributed Computing (Euro-Par) 2008<\/a>: <a href=\"http:\/\/scilytics.com\/vhpc\" target=\"scilytics.com\/vhpc\">3rd Workshop on Virtualization in High-Performance Cluster and Grid Computing (VHPC) 2008<\/a><\/i>, pages 63-71, Las Palmas de Gran Canaria, Spain, August 26-29, 2008. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-00954-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-00955-6\" target=\"publication\">10.1007\/978-3-642-00955-6<\/a>. <a href=\"javascript:showAbstract('Virtualization technology has been gaining acceptance in the scientific community due to its overall flexibility in running HPC applications. It has been reported that a specific class of applications is better suited to a particular type of virtualization scheme or implementation. For example, Xen has been shown to perform with little overhead for compute-bound applications. Such a study, although useful, does not allow us to generalize conclusions beyond the performance analysis of that application which is explicitly executed. An explanation of why the generalization described above is difficult, may be due to the versatility in applications, which leads to different overheads in virtual environments. For example, two similar applications may spend disproportionate amount of time in their respective library code when run in virtual environments. In this paper, we aim to study such potential causes by investigating the behavior and identifying patterns of various overheads for HPC benchmark applications. Based on the investigation of the overhead profiles for different benchmarks, we aim to address questions such as: Are the overhead profiles for a particular type of benchmarks (such as compute-bound) similar or are there grounds to conclude otherwise?');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tikotekar08analysis.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tikotekar08analysis.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tikotekar08analysis\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Symmetric Active\/Active High Availability for High-Performance Computing System Services: Accomplishments and Limitations<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ens-lyon.fr\/LIP\/RESO\/ccgrid2008\" target=\"www.ens-lyon.fr\/LIP\/RESO\/ccgrid2008\">8th IEEE International Symposium on Cluster Computing and the Grid (CCGrid) 2008<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2008\" target=\"xcr.cenit.latech.edu\/resilience2008\">Workshop on Resiliency in High Performance Computing (Resilience) 2008<\/a><\/i>, pages 813-818, Lyon, France, May 19-22, 2008. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3156-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CCGRID.2008.78\" target=\"publication\">10.1109\/CCGRID.2008.78<\/a>. <a href=\"javascript:showAbstract('This paper summarizes our efforts over the last 3-4 years in providing symmetric active\/active high availability for high-performance computing (HPC) system services. This work paves the way for high-level reliability, availability and serviceability in extreme-scale HPC systems by focusing on the most critical components, head and service nodes, and by reinforcing them with appropriate high availability solutions. This paper presents our accomplishments in the form of concepts and respective prototypes, discusses existing limitations, outlines possible future work, and describes the relevance of this research to other, planned efforts.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08symmetric2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann08symmetric2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08symmetric2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xin Chen, Benjamin Eckart, Xubin (Ben) He, Christian Engelmann, and Stephen L. Scott. <b>An Online Controller Towards Self-Adaptive File System Availability and Performance<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2008\" target=\"xcr.cenit.latech.edu\/hapcw2008\">5th High Availability and Performance Workshop (HAPCW) 2008<\/a>, in conjunction with the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a><\/i>, Denver, CO, USA, April 3-4, 2008. <a href=\"javascript:showAbstract('At the present time, it can be a significant challenge to build a large-scale distributed file system that simultaneously maintains both high availability and high performance. Although many fault tolerance technologies have been proposed and used in both commercial and academic distributed file systems to achieve high availability, most of them typically sacrifice performance for higher system availability. Additionally, recent studies show that system availability and performance are related to the system workload. In this paper, we analyze the correlations among availability, performance, and workloads based on a replication strategy, and we discuss the trade off between availability and performance with different workloads. Our analysis leads to the design of an online controller that can dynamically achieve optimal performance and availability by tuning the system replication policy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/chen08online.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/chen08online.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#chen08online\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Anand Tikotekar, Geoffroy Vall&eacute;e, Thomas Naughton, Hong H. Ong, Christian Engelmann, Stephen L. Scott, and Anthony M. Filippi. <b>Effects of Virtualization on a Scientific Application &#8211; Running a Hyperspectral Radiative Transfer Code on Virtual Machines<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.csm.ornl.gov\/srt\/hpcvirt08\" target=\"www.csm.ornl.gov\/srt\/hpcvirt08\">2nd Workshop on System-level Virtualization for High Performance Computing (HPCVirt) 2008<\/a>, in conjunction with the <a href=\"http:\/\/www.eurosys.org\/2008\" target=\"www.eurosys.org\/2008\">3rd ACM SIGOPS European Conference on Computer Systems (EuroSys) 2008<\/a><\/i>, pages 16-23, Glasgow, UK, March 31, 2008. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-60558-120-0. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1435452.1435455\" target=\"publication\">10.1145\/1435452.1435455<\/a>. <a href=\"javascript:showAbstract('The topic of system-level virtualization has recently begun to receive interest for high performance computing (HPC). This is in part due to the isolation and encapsulation offered by the virtual machine. These traits enable applications to customize their environments and maintain consistent software configurations in their virtual domains. Additionally, there are mechanisms that can be used for fault tolerance like live virtual machine migration. Given these attractive benefits to virtualization, a fundamental question arises, how does this effect my scientific application? We use this as the premise for our paper and observe a real-world scientific code running on a Xen virtual machine. We studied the effects of running a radiative transfer simulation, Hydrolight, on a virtual machine. We discuss our methodology and report observations regarding the usage of virtualization with this application.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tikotekar08effects.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tikotekar08effects.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tikotekar08effects\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Hong H. Ong, and Stephen L. Scott. <b>Middleware in Modern High Performance Computing System Architectures<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.iccs-meeting.org\/iccs2007\" target=\"www.iccs-meeting.org\/iccs2007\">7th International Conference on Computational Science (ICCS) 2007<\/a>, Part II: <a href=\"http:\/\/www.gup.uni-linz.ac.at\/cce2007\" target=\"www.gup.uni-linz.ac.at\/cce2007\">4th Special Session on Collaborative and Cooperative Environments (CCE) 2007<\/a><\/i>, pages 784-791, Beijing, China, May 27-30, 2007. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 3-5407-2585-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-540-72586-2_111\" target=\"publication\">10.1007\/978-3-540-72586-2_111<\/a>. <a href=\"javascript:showAbstract('A recent trend in modern high performance computing (HPC) system architectures employs lean compute nodes running a lightweight operating system (OS). Certain parts of the OS a well as other system software services are moved to service nodes in order to increase performance and scalability. This paper examines the impact of this HPC system architecture trend on HPC middleware software solutions, which traditionally equip HPC systems with advanced features, such as parallel and distributed programming models, appropriate system resource management mechanisms, remote application steering and user interaction techniques. Since the approach of keeping the compute node software stack small and simple is orthogonal to the middleware concept of adding missing OS features between OS and application, the role and architecture of middleware in modern HPC systems needs to be revisited. The result is a paradigm shift in HPC middleware design, where single middleware services are moved to service nodes, while runtime environments (RTEs) continue to reside on compute nodes.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07middleware.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07middleware.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07middleware\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Transparent Symmetric Active\/Active Replication for Service-Level High Availability<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ccgrid07.lncc.br\" target=\"ccgrid07.lncc.br\">7th IEEE International Symposium on Cluster Computing and the Grid (CCGrid) 2007<\/a>: <a href=\"http:\/\/www.lri.fr\/ fedak\/gp2pc-07\" target=\"www.lri.fr\/ fedak\/gp2pc-07\">7th International Workshop on Global and Peer-to-Peer Computing (GP2PC) 2007<\/a><\/i>, pages 755-760, Rio de Janeiro, Brazil, May 14-17, 2007. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2833-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CCGRID.2007.116\" target=\"publication\">10.1109\/CCGRID.2007.116<\/a>. <a href=\"javascript:showAbstract('As service-oriented architectures become more important in parallel and distributed computing systems, individual service instance reliability as well as appropriate service redundancy becomes an essential necessity in order to increase overall system availability. This paper focuses on providing redundancy strategies using service-level replication techniques. Based on previous research using symmetric active\/active replication, this paper proposes a transparent symmetric active\/active replication approach that allows for more reuse of code between individual service-level replication implementations by using a virtual communication layer. Service- and client-side interceptors are utilized in order to provide total transparency. Clients and servers are unaware of the replication infrastructure as it provides all necessary mechanisms internally.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07transparent.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07transparent.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07transparent\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Hong H. Ong, Geoffroy R. Vall&eacute;e, and Thomas Naughton. <b>Configurable Virtualized System Environments for High Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.csm.ornl.gov\/srt\/hpcvirt07\" target=\"www.csm.ornl.gov\/srt\/hpcvirt07\">1st Workshop on System-level Virtualization for High Performance Computing (HPCVirt) 2007<\/a>, in conjunction with the <a href=\"http:\/\/www.eurosys.org\/2008\" target=\"www.eurosys.org\/2008\">2nd ACM SIGOPS European Conference on Computer Systems (EuroSys) 2007<\/a><\/i>, Lisbon, Portugal, March 20, 2007. <a href=\"javascript:showAbstract('Existing challenges for current terascale high performance computing (HPC) systems are increasingly hampering the development and deployment efforts of system software and scientific applications for next-generation petascale systems. The expected rapid system upgrade interval toward petascale scientific computing demands an incremental strategy for the development and deployment of legacy and new large-scale scientific applications that avoids excessive porting. Furthermore, system software developers as well as scientific application developers require access to large-scale testbed environments in order to test individual solutions at scale. This paper proposes to address these issues at the system software level through the development of a virtualized system environment (VSE) for scientific computing. The proposed VSE approach enables plug-and-play supercomputing through desktop-to-cluster-to-petaflop computer system-level virtualization based on recent advances in hypervisor virtualization technologies. This paper describes the VSE system architecture in detail, discusses needed tools for VSE system management and configuration, and presents respective VSE use case scenarios.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07configurable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07configurable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07configurable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Towards High Availability for High-Performance Computing System Services: Accomplishments and Limitations<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2006\" target=\"xcr.cenit.latech.edu\/hapcw2006\">4th High Availability and Performance Workshop (HAPCW) 2006<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.krellinst.org\" target=\"lacsi.krellinst.org\">7th Los Alamos Computer Science Institute (LACSI) Symposium 2006<\/a><\/i>, Santa Fe, NM, USA, October 17, 2006. <a href=\"javascript:showAbstract('During the last several years, our teams at Oak Ridge National Laboratory, Louisiana Tech University, and Tennessee Technological University focused on efficient redundancy strategies for head and service nodes of high-performance computing (HPC) systems in order to pave the way for high availability (HA) in HPC. These nodes typically run critical HPC system services, like job and resource management, and represent single points of failure and control for an entire HPC system. The overarching goal of our research is to provide high-level reliability, availability, and serviceability (RAS) for HPC systems by combining HA and HPC technology. This paper summarizes our accomplishments, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06towards.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann06towards.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Li Ou, Xin Chen, Xubin (Ben) He, Christian Engelmann, and Stephen L. Scott. <b>Achieving Computational I\/O Effciency in a High Performance Cluster Using Multicore Processors<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2006\" target=\"xcr.cenit.latech.edu\/hapcw2006\">4th High Availability and Performance Workshop (HAPCW) 2006<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.krellinst.org\" target=\"lacsi.krellinst.org\">7th Los Alamos Computer Science Institute (LACSI) Symposium 2006<\/a><\/i>, Santa Fe, NM, USA, October 17, 2006. <a href=\"javascript:showAbstract('Cluster computing has become one of the most popular platforms for high-performance computing today. The recent popularity of multicore processors provides a flexible way to increase the computational capability of clusters. Although the system performance may improve with multicore processors in a cluster, I\/O requests initiated by multiple cores may saturate the I\/O bus, and furthermore increase the latency by issuing  multiple non-contiguous disk accesses. In this paper, we propose an asymmetric collective I\/O for multicore processors to improve multiple non-contiguous accesses. In our configuration, one core in each multicore processor is designated as the coordinator, and others serve as computing cores. The coordinator is responsible for aggregating I\/O operations from computing cores and submitting a contiguous request. The coordinator allocates contiguous memory buffers on behalf of other cores to avoid redundant data copies.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ou06achieving.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ou06achieving.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ou06achieving\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>RMIX: A Dynamic, Heterogeneous, Reconfigurable Communication Framework<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.iccs-meeting.org\/iccs2006\" target=\"www.iccs-meeting.org\/iccs2006\">6th International Conference on Computational Science (ICCS) 2006<\/a>, Part II: <a href=\"http:\/\/www.gup.uni-linz.ac.at\/cce2006\" target=\"www.gup.uni-linz.ac.at\/cce2006\">3rd Special Session on Collaborative and Cooperative Environments (CCE) 2006<\/a><\/i>, pages 573-580, Reading, UK, May 28-31, 2006. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 3-540-34381-4. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/11758525_77\" target=\"publication\">10.1007\/11758525_77<\/a>. <a href=\"javascript:showAbstract('RMIX is a dynamic, heterogeneous, reconfigurable communication framework that allows software components to communicate using various RMI\/RPC protocols, such as ONC RPC, Java RMI and SOAP, by facilitating dynamically loadable provider plug-ins to supply different protocol stacks. With this paper, we present a native (C-based), flexible, adaptable, multi-protocol RMI\/RPC communication framework that complements the Java-based RMIX variant previously developed by our partner team at Emory University. Our approach offers the same multi-protocol RMI\/RPC services and advanced invocation semantics via a C-based interface that does not require an object-oriented programming language. This paper provides a detailed description of our RMIX framework architecture and some of its features. It describes the general use case of the RMIX framework and its integration into the Harness metacomputing environment in the form of a plug-in.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06rmix.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann06rmix.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06rmix\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Active\/Active Replication for Highly Available HPC System Services<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ares-conference.eu\/ares2006\" target=\"www.ares-conference.eu\/ares2006\">1st International Conference on Availability, Reliability and Security (ARES) 2006<\/a>: 1st International Workshop on Frontiers in Availability, Reliability and Security (FARES) 2006<\/i>, pages 639-645, Vienna, Austria, April 20-22, 2006. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2567-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ARES.2006.23\" target=\"publication\">10.1109\/ARES.2006.23<\/a>. <a href=\"javascript:showAbstract('Today`s high performance computing systems have several reliability deficiencies resulting in availability and serviceability issues. Head and service nodes represent a single point of failure and control for an entire system as they render it inaccessible and unmanageable in case of a failure until repair, causing a significant downtime. This paper introduces two distinct replication methods (internal and external) for providing symmetric active\/active high availability for multiple head and service nodes running in virtual synchrony. It presents a comparison of both methods in terms of expected correctness, ease-of-use and performance based on early results from ongoing work in providing symmetric active\/active high availability for two HPC system services (TORQUE and PVFS metadata server). It continues with a short description of a distributed mutual exclusion algorithm and a brief statement regarding the handling of Byzantine failures. This paper concludes with an overview of past and ongoing work, and a short summary of the presented research.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06active.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann06active.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06active\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Stephen L. Scott. <b>Concepts for High Availability in Scientific High-End Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2005\" target=\"xcr.cenit.latech.edu\/hapcw2005\">3rd High Availability and Performance Workshop (HAPCW) 2005<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.rice.edu\/symposium\/agenda_2005\" target=\"lacsi.rice.edu\/symposium\/agenda_2005\">6th Los Alamos Computer Science Institute (LACSI) Symposium 2005<\/a><\/i>, Santa Fe, NM, USA, October 11, 2005. <a href=\"javascript:showAbstract('Scientific high-end computing (HEC) has become an important tool for scientists world-wide to understand problems, such as in nuclear fusion, human genomics and nanotechnology. Every year, new HEC systems emerge on the market with better performance and higher scale. With only very few exceptions, the overall availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically due to the recent trend towards capability computing. In this paper, we analyze the existing deficiencies of current HEC systems and present several high availability concepts to counter the experienced loss of availability and to alleviate the expected impact on next-generation systems. We explain the application of these concepts to current and future HEC systems and list past and ongoing related research. This paper closes with a short summary of the presented work and a brief discussion of future efforts.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05concepts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05concepts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05concepts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Stephen L. Scott. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/coset.irisa.fr\" target=\"coset.irisa.fr\">2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2) 2005<\/a>, in conjunction with the <a href=\"http:\/\/ics05.csail.mit.edu\" target=\"ics05.csail.mit.edu\">19th ACM International Conference on Supercomputing (ICS) 2005<\/a><\/i>, Cambridge, MA, USA, June 19, 2005. <a href=\"javascript:showAbstract('Ultra-scale architectures for scientific high-end computing with tens to hundreds of thousands of processors, such as the IBM Blue Gene\/L and the Cray X1, suffer from availability deficiencies, which impact the efficiency of running computational jobs by forcing frequent checkpointing of applications. Most systems are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services, such as the job scheduler or MPI, or even of the entire machine. In this paper, we present a flexible, pluggable and component-based high availability framework that expands today`s effort in high availability computing of keeping a single server alive to include all machines cooperating in a high-end scientific computing environment, while allowing adaptation to system properties and application needs.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chokchai (Box) Leangsuksun, Venkata K. Munganuru, Tong Liu, Stephen L. Scott, and Christian Engelmann. <b>Asymmetric Active-Active High Availability for High-end Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/coset.irisa.fr\" target=\"coset.irisa.fr\">2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2) 2005<\/a>, in conjunction with the <a href=\"http:\/\/ics05.csail.mit.edu\" target=\"ics05.csail.mit.edu\">19th ACM International Conference on Supercomputing (ICS) 2005<\/a><\/i>, Cambridge, MA, USA, June 19, 2005. <a href=\"javascript:showAbstract('Linux clusters have become very popular for scientific computing at research institutions world-wide, because they can be easily deployed at a fairly low cost. However, the most pressing issues of today`s cluster solutions are availability and serviceability. The conventional Beowulf cluster architecture has a single head node connected to a group of compute nodes. This head node is a typical single point of failure and control, which severely limits availability and serviceability by effectively cutting off healthy compute nodes from the outside world upon overload or failure. In this paper, we describe a paradigm that addresses this issue using asymmetric active-active high availability. Our framework comprises of n + 1 head nodes, where n head nodes are active in the sense that they provide services to simultaneously incoming user requests. One standby server monitors all active servers and performs a fail-over in case of a detected outage. We present a prototype implementation based on a 2 + 1 solution and discuss initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/leangsuksun05asymmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/leangsuksun05asymmetric.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#leangsuksun05asymmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>A Lightweight Kernel for the Harness Metacomputing Framework<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ipdps.org\/ipdps2005\" target=\"www.ipdps.org\/ipdps2005\">19th IEEE International Parallel and Distributed Processing Symposium (IPDPS) 2005<\/a>: <a href=\"http:\/\/www.cs.umass.edu\/ rsnbrg\/hcw2005\" target=\"www.cs.umass.edu\/ rsnbrg\/hcw2005\">14th Heterogeneous Computing Workshop (HCW) 2005<\/a><\/i>, Denver, CO, USA, April 4, 2005. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2312-9. ISSN 1530-2075. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/IPDPS.2005.34\" target=\"publication\">10.1109\/IPDPS.2005.34<\/a>. <a href=\"javascript:showAbstract('Harness is a pluggable heterogeneous Distributed Virtual Machine (DVM) environment for parallel and distributed scientific computing. This paper describes recent improvements in the Harness kernel design. By using a lightweight approach and moving previously integrated system services into software modules, the software becomes more versatile and adaptable. This paper outlines these changes and explains the major Harness kernel components in more detail. A short overview is given of ongoing efforts in integrating RMIX, a dynamic heterogeneous reconfigurable communication framework, into the Harness environment as a new plug-in software module. We describe the overall impact of these changes and how they relate to other ongoing work.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05lightweight.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05lightweight.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05lightweight\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, and George A. (Al) Geist. <b>High Availability through Distributed Control<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2004\" target=\"xcr.cenit.latech.edu\/hapcw2004\">2nd High Availability and Performance Workshop (HAPCW) 2004<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.rice.edu\/symposium\/agenda_2004\" target=\"lacsi.rice.edu\/symposium\/agenda_2004\">5th Los Alamos Computer Science Institute (LACSI) Symposium 2004<\/a><\/i>, Santa Fe, NM, USA, October 12, 2004. <a href=\"javascript:showAbstract('Cost-effective, flexible and efficient scientific simulations in cutting-edge research areas utilize huge high-end computing resources with thousands of processors. In the next five to ten years the number of processors in such computer systems will rise to tens of thousands, while scientific application running times are expected to increase further beyond the Mean-Time-To-Interrupt (MTTI) of hardware and system software components. This paper describes the ongoing research in heterogeneous adaptable reconfigurable networked systems (Harness) and its recent achievements in the area of high availability distributed virtual machine environments for parallel and distributed scientific computing. It shows how a distributed control algorithm is able to steer a distributed virtual machine process in virtual synchrony while maintaining consistent replication for high availability. It briefly illustrates ongoing work in heterogeneous reconfigurable communication frameworks and security mechanisms. The paper continues with a short overview of similar research in reliable group communication frameworks, fault-tolerant process groups and highly available distributed virtual processes. It closes with a brief discussion of possible future research directions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann04high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann04high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann04high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xubin (Ben) He, Li Ou, Stephen L. Scott, and Christian Engelmann. <b>A Highly Available Cluster Storage System using Scavenging<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2004\" target=\"xcr.cenit.latech.edu\/hapcw2004\">2nd High Availability and Performance Workshop (HAPCW) 2004<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.rice.edu\/symposium\/agenda_2004\" target=\"lacsi.rice.edu\/symposium\/agenda_2004\">5th Los Alamos Computer Science Institute (LACSI) Symposium 2004<\/a><\/i>, Santa Fe, NM, USA, October 12, 2004. <a href=\"javascript:showAbstract('Highly available data storage for high-performance computing is becoming increasingly more critical as high-end computing systems scale up in size and storage systems are developed around network-centered architectures. A promising solution is to harness the collective storage potential of individual workstations much as we harness idle CPU cycles due to the excellent price\/performance ratio and low storage usage of most commodity workstations. For such a storage system, metadata consistency is a key issue assuring storage system availability as well as data reliability. In this paper, we present a decentralized metadata management scheme that improves storage availability without sacrificing performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/he04highly.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/he04highly.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#he04highly\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>A Diskless Checkpointing Algorithm for Super-scale Architectures Applied to the Fast Fourier Transform<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.cs.msstate.edu\/ clade2003\" target=\"www.cs.msstate.edu\/ clade2003\">Challenges of Large Applications in Distributed Environments Workshop (CLADE) 2003<\/a>, in conjunction with the <a href=\"http:\/\/csag.ucsd.edu\/HPDC-12\" target=\"csag.ucsd.edu\/HPDC-12\">12th IEEE International Symposium on High Performance Distributed Computing (HPDC) 2003<\/a><\/i>, pages 47, Seattle, WA, USA, June 21, 2003. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-1984-9. DOI <a href=\"http:\/\/dx.doi.org\/xpls\/abs_all.jsp?arnumber=4159902\" target=\"publication\">xpls\/abs_all.jsp?arnumber=4159902<\/a>. <a href=\"javascript:showAbstract('This paper discusses the issue of fault-tolerance in distributed computer systems with tens or hundreds of thousands of diskless processor units. Such systems, like the IBM Blue Gene\/L, are predicted to be deployed in the next five to ten years. Since a 100,000-processor system is going to be less reliable, scientific applications need to be able to recover from occurring failures more efficiently. In this paper, we adapt the present technique of diskless checkpointing to such huge distributed systems in order to equip existing scientific algorithms with super-scalable fault-tolerance. First, we discuss the method of diskless checkpointing, then we adapt this technique to super-scale architectures and finally we present results from an implementation of the Fast Fourier Transform that uses the adapted technique to achieve super-scale fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann03diskless.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann03diskless.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann03diskless\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, and George A. (Al) Geist. <b>Distributed Peer-to-Peer Control in Harness<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.science.uva.nl\/events\/ICCS2002\" target=\"www.science.uva.nl\/events\/ICCS2002\">2nd International Conference on Computational Science (ICCS) 2002<\/a>, Part II: Workshop on Global and Collaborative Computing<\/i>, pages 720-727, Amsterdam, The Netherlands, April 21-24, 2002. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 3-540-43593-X. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/content\/l537ujfwt8yta2dp\" target=\"publication\">content\/l537ujfwt8yta2dp<\/a>. <a href=\"javascript:showAbstract('Harness is an adaptable fault-tolerant virtual machine environment for next-generation heterogeneous distributed computing developed as a follow on to PVM. It additionally enables the assembly of applications from plug-ins and provides fault-tolerance. This work describes the distributed control, which manages global state replication to ensure a high-availability of service. Group communication services achieve an agreement on an initial global state and a linear history of global state changes at all members of the distributed virtual machine. This global state is replicated to all members to easily recover from single, multiple and cascaded faults. A peer-to-peer ring network architecture and tunable multi-point failure conditions provide heterogeneity and scalability. Finally, the integration of the distributed control into the multi-threaded kernel architecture of Harness offers a fault-tolerant global state database service for plug-ins and applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann02distributed.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann02distributed.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann02distributed\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"posters\"><\/a>Peer-reviewed Conference Posters<\/h4>\n<ol>\n<li>Swen Boehm, Craig A. Bridges, Patrick Widener, Terry Jones, Sheikh Ghafoor, Christian Engelmann, and Olga Kuchar. <b>From Automated Experiments and Simulations to Reusable Scientific Evidence<\/b>. Poster at the <a href=\"http:\/\/www.montereydataconference.org\" target=\"www.montereydataconference.org\">Monterey Data Conference<\/a>, Monterey, CA, USA, August 24-26, 2026. <a href=\"javascript:showAbstract('Scientific research increasingly depends on automated platforms and heterogeneous instruments that generate rich, multi-modal data. Yet fragmentation across domain-specific systems, proprietary formats, and siloed repositories impedes data discovery, integration, and reuse. Without systematic semantic representation and comprehensive provenance, high-quality experimental data loses much of its scientific value. The INTERSECT Scientific data layer (SDL) provides an integrated, ontology-driven ecosystem that connects scientific platforms, workflows, and data management services into a coherent whole. Built on a system-of-systems architecture and grounded in Linked Data Platform (LDP) principles, the SDL enables modular integration of diverse services while preserving interoperability across scientific domains.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm26automated.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#boehm26automated\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen Boehm, Terry Jones, Patrick Widener, Christian Engelman, and Olga Kuchar. <b>INTERSECT Scientific Data Layer: A Federated, Modular Framework for Scientific Data Management<\/b>. Poster at the <a href=\"http:\/\/data-science.llnl.gov\/d3\" target=\"data-science.llnl.gov\/d3\">Department of Energy (DOE) Data Days (D3) Workshop<\/a>, Chantilly, VA, USA, March 3-5, 2026. <a href=\"javascript:showAbstract('Scientific research increasingly depends on automated platforms and heterogeneous instruments that generate rich, multi-modal data. Yet fragmentation across domain-specific systems, proprietary formats, and siloed repositories impedes data discovery, integration, and reuse. Without systematic semantic representation and comprehensive provenance, high-quality experimental data loses much of its scientific value.  The INTERSECT Scientific data layer (SDL) provides an integrated, ontology-driven ecosystem that connects scientific platforms, workflows, and data management services into a coherent whole. Built on a system-of-systems architecture and grounded in Linked Data Platform (LDP) principles, the SDL enables modular integration of diverse services while preserving interoperability across scientific domains. Core ontologies such as SSN\/SOSA for sensor and observation modeling, DCAT for resource cataloging, and PROV-O for provenance tracking provide a semantic backbone that ensures all entities - data, instruments, workflows, and results - are described in a machine-actionable, reusable way.  The SDL offers semantic-first design, a microservices foundation, separation of concerns, and content negotiation:  - Semantic-First Design: RDF is the native data model, not   an auxiliary export format. Semantic richness is preserved   throughout the data lifecycle, from instrumental observations   through processing pipelines to publication, enabling FAIR   data by design. - Microservices Foundation: Modular, independently deployable   services (Catalog Service, Storage Service, Repository   Service, Registry Service) coordinate through shared   semantic libraries and standard ontologies, solving the   distributed consistency challenge inherent in semantic   systems. - Separation of Concerns: Semantic metadata (RDF triples in   triple stores) is decoupled from data artifacts (files in   object storage) with URIs providing semantic linking. This   enables independent scaling of metadata management and   storage infrastructure while maintaining coherent provenance   relationships. - Content Negotiation: Services accept and return data in   multiple RDF serializations (Turtle, JSON-LD, RDF\/XML) and   domain-specific formats (CSV, HDF5, instrument formats),   supporting diverse tools and workflows while maintaining    semantic consistency.  The SDL natively implements FAIR principles through semantic-first architecture. Persistent URIs and SPARQL endpoints enable discovery via machine-readable metadata (findable). Standard HTTP protocols and LDP containers support predictable REST-like access patterns (accessible). Composed W3C ontologies ensure semantic compatibility across domains (interoperable). Comprehensive end-to-end provenance and structured metadata make datasets suitable for both human researchers and AI systems (reusable).');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm26intersect.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#boehm26intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Swen Boehm, Michael Brim, Jack Lange, Thomas Naughton, Patrick Widener, Ben Mintz, and Rohit Srivastava. <b>INTERSECT: The Open Federated Architecture for the Laboratory of the Future<\/b>. Poster at the <a href=\"http:\/\/icpp23.sci.utah.edu\/\" target=\"icpp23.sci.utah.edu\/\">52nd International Conference on Parallel Processing (ICPP) 2023<\/a>, Salt Lake City, UT, USA, August 7-10, 2023. <a href=\"javascript:showAbstract('The open Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) architecture connects scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing center to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery and evaluation. Its a novel approach consists of science use case design patterns, a system of systems architecture, and a microservice architecture.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23intersect.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann23intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Mohit Kumar. <b>Resilience Design Patterns: A Structured Modeling Approach of Resilience in Computing Systems<\/b>. Poster at the <a href=\"http:\/\/www.bnl.gov\/modsim2022\" target=\"www.bnl.gov\/modsim2022\">Workshop on Modeling and Simulation of Systems and Applications (ModSim) 2022<\/a>, Seattle, WA, USA, August 10-12, 2022. <a href=\"javascript:showAbstract('Resilience to faults, errors, and failures in extreme-scale high-performance computing (HPC) systems is a critical challenge. Resilience design patterns (Figure 1) offer a new, structured hardware\/software design approach for improving resilience by identifying and evaluating repeatedly occurring resilience problems and coordinating corresponding solutions. Initial work identified and formalized these patterns and developed a proof-of-concept prototype to demonstrate portable resilience. This recent work created performance, reliability, and availability models for each of the identified 15 structural resilience design patterns and a modeling tool that allows (1) exploring the performance, reliability, and availability of each pattern, and (2) investigating the trade-offs be-tween patterns and pattern combinations.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22resilience.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann22resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Rizwan Ashraf, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>Real-Time Assessment of Supercomputer Status by a Comprehensive Informative Metric through Streaming Processing<\/b>. Poster at the  <a href=\"http:\/\/cci.drexel.edu\/bigdata\/bigdata2018\" target=\"cci.drexel.edu\/bigdata\/bigdata2018\">6th IEEE International Conference on Big Data (BigData) 2018<\/a>,  Seattle, WA, USA, December 10-13, 2018. <a href=\"javascript:showAbstract('Supercomputers are complex systems used to simulate, understand and solve real-world problems. In order to operate these systems efficiently and for the purpose of their maintainability, an accurate, concise, and timely determination of system status is crucial for its users and operators. However, this determination is challenging due to intricately connected heterogeneous software and hardware components, and due to sheer scale of such machines. In this poster, we demonstrate work-in-progress towards realization of a real-time monitoring framework for the 18,688-node Titan supercomputer at Oak Ridge Leadership Computing Facility (OLCF). Toward this end, we discuss the use of metrics which present a one-dimensional view of the system generating various types of information from 1000s of components and utilization statistics from 100s of user applications in near real-time. We demonstrate the efficacy of these metrics to understand and visualize raw log data generated by the system which otherwise may compose of 1000s of dimensions. We also demonstrate the architecture of proposed real-time stream processing framework which integrates, processes, analyzes, visualizes and stores system log data from an array of system components..');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18realtime.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hui18realtime\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>A Comprehensive Informative Metric for Summarizing HPC System Status<\/b>. Poster at the <a href=\"http:\/\/ldav.org\" target=\"ldav.org\">8th IEEE Symposium on Large Data Analysis and   Visualization<\/a> in conjunction with the   <a href=\"http:\/\/ieeevis.org\/year\/2018\" target=\"ieeevis.org\/year\/2018\">8th IEEE Vis 2018<\/a>,  Berlin, Germany, October 21, 2018. <a href=\"javascript:showAbstract('It remains a major challenge to effectively summarize and visualize in a comprehensive form the status of a complex computer system, such as the Titan supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). In the ongoing research highlighted in this poster, we present system information entropy (SIE), a newly developed system metric that leverages the powers of traditional machine learning techniques and information theory. By compressing the multi-variant multi-dimensional event information recorded during the operation of the targeted system into a single time series of SIE, we demonstrate that the historical system status can be sensitively summarized in form of SIE and visualized concisely and comprehensively.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18comprehensive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hui18comprehensive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Rizwan Ashraf. <b>Modeling and Simulation of Extreme-Scale Systems for Resilience by Design<\/b>. Poster at the <a href=\"http:\/\/www.bnl.gov\/modsim2018\" target=\"www.bnl.gov\/modsim2018\">Workshop on Modeling and Simulation of Systems and Applications<\/a>, Seattle, WA, USA, August 15-17, 2018. <a href=\"javascript:showAbstract('Resilience is a serious concern for extreme-scale high-performance computing (HPC). While the HPC community has developed various resilience solutions, the solution space remains fragmented. We created a structured approach to the design, evaluation and optimization of HPC resilience using the concept of design patterns. A design pattern describes a generalized solution to a repeatedly occurring problem. We identified the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. Each well-known solution that addresses a specific resilience challenge is described in the form of a design pattern. We developed a resilience design pattern specification, language and catalog, which can be used by system architects, system software and library developers, application programmers, as well as users and operators as essential building blocks when designing and deploying resilience solutions. The resilience design pattern approach provides a unique opportunity for design space exploration. As each resilience solution is abstracted as a pattern and each solution&amp;#39;s properties are defined by pattern parameters, vertical and horizontal pattern compositions can describe the resilience capabilities of an entire HPC system. This permits the investigation of beneficial or counterproductive interactions between patterns and of the performance, resilience, and power consumption trade-off between different pattern parameters and compositions. The ultimate goal is to make resilience an integral part of the HPC hardware\/software ecosystem by coordinating the various existing resilience solutions in a design space exploration process, such that the burden for providing resilience is on the system by design and not on the user as an afterthought. We are in the early stages of developing a novel design space exploration tool that enables this investigation using modeling and simulation. We developed performance and resilience models for each resilience design pattern. We also leverage results from the Catalog project, a collaborative effort between Oak Ridge National Laboratory, Argonne National Laboratory and Lawrence Livermore National Laboratory that developed models of the faults, errors and failures in today's HPC systems. We also leverage recent results from the same project by Lawrence Livermore National Laboratory in application reliability patterns. The planned research extends and combines this work to model the performance, resilience, and power consumption of an entire HPC system, initially at node-level granularity, and to simulate the dynamic interactions between deployed resilience solutions and the rest of the system. In the next iteration, finer-grain modeling and simulation, such as at the computational unit level, is used to increase accuracy. This work leverages the experience of the investigators in parallel discrete event simulation of extreme-scale systems, such as the Extreme-scale Simulator (xSim). The current state of the art in resilience modeling and simulation is fragmented as well. There is currently no such design space exploration tool. Instead, each resilience solution is typically investigated separately. There is only a small amount of work on multi-resilience solutions, including by the investigators. While there is work in investigating the performance\/resilience trade-off space, there is almost no work in including power consumption.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18modeling2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann18modeling2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Onkar Patil, Saurabh Hukerikar, Frank Mueller, and Christian Engelmann. <b>Exploring Use Cases for Non-Volatile Memories in Support of HPC Resilience<\/b>. Poster at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">30th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2017<\/a>, Denver, CO, USA, November 12-17, 2017. <a href=\"javascript:showAbstract('Improving resilience and creating resilient architectures is one of the major goals of exascale computing. With the advent of Non-volatile memory technologies, memory architectures with persistent memory regions will be a significant part of future architectures. There is potential to use them in more than one way to benefit different applications. We look to take advantage of this technology to enable more fine-grained and novel methodology that will improve resilience and efficiency of exascale applications. We have developed three modes of memory usage for persistent memory to enable efficient checkpointing in HPC applications. We have developed a simple API that is evaluated with the DGEMM benchmark on a 16-node cluster with independent SSDs on every node. Our aim is to build on this work and enable static and dynamic runtime systems that will inherently make the HPC applications more fault-tolerant and resistant to errors.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/patil17exploring.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#patil17exploring\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Frank Mueller, Christian Engelmann, Rolf Riesen, and Kurt Ferreira. <b>Detection and Correction of Silent Data Corruption for Large-Scale High-Performance Computing<\/b>. Poster at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">24th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2011<\/a>, Seattle, WA, USA, November 12-18, 2011. <a href=\"javascript:showAbstract('Faults have become the norm rather than the exception for high-end computing on clusters with 10s\/100s of thousands of cores. Exacerbating this situation, some of these faults will not be detected, manifesting themselves as silent errors that will corrupt memory while applications continue to operate and report incorrect results. This poster introduces RedMPI, an MPI library which resides in the MPI profiling layer. RedMPI is capable of both online detection and correction of soft errors that occur in MPI applications without requiring any modifications to the application source. By providing redundancy, RedMPI is capable of transparently detecting corrupt messages from MPI processes that become faulted during execution. Furthermore, with triple redundancy RedMPI additionally &amp;#34;votes&amp;#34; out MPI messages of a faulted process by replacing corrupted results with corrected results from unfaulted processes. We present an experimental evaluation of RedMPI on an assortment of applications to demonstrate the effectiveness of this approach.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#fiala11detection\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Kurt Ferreira, Frank Mueller, and Christian Engelmann. <b>A Tunable, Software-based DRAM Error Detection and Correction Library for HPC<\/b>. Poster at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">24th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2011<\/a>, Seattle, WA, USA, November 12-18, 2011. <a href=\"javascript:showAbstract('Proposed exascale systems will present a number of considerable resiliency challenges. In particular, DRAM soft-errors, or bit-flips, are expected to greatly increase due to the increased memory density of these systems. Current hardware-based fault-tolerance methods will be unsuitable for addressing the expected soft error frequency rate. As a result, additional software will be needed to address this challenge. In this paper we introduce LIBSDC, a tunable, transparent silent data corruption detection and correction library for HPC applications. LIBSDC provides comprehensive SDC protection for program memory by implementing on-demand page integrity verification by utilizing the MMU. Experimental  benchmarks with Mantevo HPCCG show that once tuned, LIBSDC is able to achieve SDC protection with less than 100% overhead of resources.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#fiala11tunable2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, George Ostrouchov, Chokchai (Box) Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller, Chao Wang, Arun B. Nagarajan, and Jyothish Varma. <b>A Tunable Holistic Resiliency Approach for High-Performance Computing Systems<\/b>. Poster at the <a href=\"http:\/\/institute.lanl.gov\/resilience\/conferences\/2009\" target=\"institute.lanl.gov\/resilience\/conferences\/2009\">National HPC Workshop on Resilience 2009<\/a>, Arlington, VA, USA, August 12-14, 2009. <a href=\"javascript:showAbstract('In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation extreme-scale high-performance computing (HPC) systems. This poster describes our past and ongoing efforts in novel fault resilience technologies for HPC. Presented work includes proactive fault resilience techniques, system and application reliability models and analyses, failure prediction, transparent process- and virtual-machine-level migration, and trade-off models for combining preemptive migration with checkpoint\/restart. This poster summarizes our work and puts all individual technologies into context with a proposed holistic fault resilience framework.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott09tunable2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott09tunable2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, Christian Engelmann, and Hong H. Ong. <b>System-level Virtualization for for High-Performance Computing<\/b>. Poster at the <a href=\"http:\/\/institute.lanl.gov\/resilience\/conferences\/2009\" target=\"institute.lanl.gov\/resilience\/conferences\/2009\">National HPC Workshop on Resilience 2009<\/a>, Arlington, VA, USA, August 12-14, 2009. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel system software solutions for providing a virtual system environment (VSE) for next-generation extreme-scale high-performance computing (HPC) systems and beyond. The poster showcases results of developed proof-of-concept implementations and performed theoretical analyses, outlines planned research and development activities, and presents respective initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott09systemlevel.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott09systemlevel\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, George Ostrouchov, Chokchai (Box) Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller, Chao Wang, Arun B. Nagarajan, and Jyothish Varma. <b>A Tunable Holistic Resiliency Approach for High-Performance Computing Systems<\/b>. Poster at the <a href=\"http:\/\/ppopp09.rice.edu\" target=\"ppopp09.rice.edu\">14th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP) 2009<\/a>, Raleigh, NC, USA, February 14-18, 2009. <a href=\"javascript:showAbstract('In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation extreme-scale high-performance computing (HPC) systems. This poster describes our past and ongoing efforts in novel fault resilience technologies for HPC. Presented work includes proactive fault resilience techniques, system and application reliability models and analyses, failure prediction, transparent process- and virtual-machine-level migration, and trade-off models for combining preemptive migration with checkpoint\/restart. This poster summarizes our work and puts all individual technologies into context with a proposed holistic fault resilience framework.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott09tunable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott09tunable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>George A. (Al) Geist, Christian Engelmann, Jack J. Dongarra, George Bosilca, Magdalena M. S&#322;awi&#324;ska, and Jaros&#322;aw K. S&#322;awi&#324;ski. <b>The Harness Workbench: Unified and Adaptive Access to Diverse High-Performance Computing Platforms<\/b>. Poster at the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a>, Denver, CO, USA, March 30 &#8211; April 5, 2008. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel software solutions for providing unified and adaptive access to diverse high-performance computing (HPC) platforms. The poster showcases developed proof-of-concept implementations of tools and mechanisms that simplify scientific application development and deployment tasks, such that only minimal adaptation is needed when moving from one HPC system to another or after HPC system upgrades.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/geist08harness.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#geist08harness\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Christian Engelmann, Hong H. Ong, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, George Ostrouchov, Chokchai (Box) Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller, Chao Wang, Arun B. Nagarajan, Jyothish Varma, Xubin (Ben) He, Li Ou, and Xin Chen. <b>Resiliency for High-Performance Computing Systems<\/b>. Poster at the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a>, Denver, CO, USA, March 30 &#8211; April 5, 2008. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel system software solutions for providing high-level reliability, availability and serviceability (RAS) for next-generation extreme-scale high-performance computing (HPC) systems and beyond. The poster showcases results of developed proof-of-concept implementations and performed theoretical analyses, outlines planned research and development activities, and presents respective initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott08resiliency.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott08resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, Christian Engelmann, and Hong H. Ong. <b>System-level Virtualization for for High-Performance Computing<\/b>. Poster at the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a>, Denver, CO, USA, March 30 &#8211; April 5, 2008. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel system software solutions for providing a virtual system environment (VSE) for next-generation extreme-scale high-performance computing (HPC) systems and beyond. The poster showcases results of developed proof-of-concept implementations and performed theoretical analyses, outlines planned research and development activities, and presents respective initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott08systemlevel.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott08systemlevel\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"whitepapers\"><\/a>White Papers<\/h4>\n<ol>\n<li>Ryan Adamson and Christian Engelmann. <b>Cybersecurity and Privacy for Instrument-to-Edge-to-Center Scientific Computing Ecosystems<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/2021ascr-cybersecurity\" target=\"www.orau.gov\/2021ascr-cybersecurity\">ASCR Workshop on Cybersecurity and Privacy for Scientific  Computing Ecosystems<\/a><\/i>, November 3-5, 2021. <a href=\"javascript:showAbstract('The DOE&amp;#39;s Artificial Intelligence (AI) for Science report outlines the need for intelligent systems, instruments, and facilities to enable science breakthroughs with autonomous experiments, 'self-driving' laboratories, smart manufacturing, and AI-driven design, discovery and evaluation. The DOE's Computational Facilities Research Workshop report identifies intelligent systems\/facilities as a challenge with enabling automation and eliminating human-in-the-loop needs as a cross-cutting theme. Autonomous experiments, 'self-driving' laboratories and smart manufacturing employ machine-in-the-loop intelligence for decision-making. Human-in-the-loop needs are reduced by an autonomous online control that collects experiment data, analyzes it, and takes appropriate operational actions in real time to steer an ongoing or plan the next experiment. DOE laboratories are currently in the process of developing and deploying federated hardware\/software architectures for connecting instruments with edge and center computing resources to autonomously collect, transfer, store, process, curate, and archive scientific data. These new instrument-to-edge-to-center scientific ecosystems face several cybersecurity and privacy challenges.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/adamson21cybersecurity.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#adamson21cybersecurity\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mingyan Li, Robert A. Bridges, Pablo Moriano, Christian Engelmann, Feiyi Wang, and Ryan Adamson. <b>Toward Effective Security\/Reliability Situational Awareness via Concurrent Security-or-Fault Analytics <\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/2021ascr-cybersecurity\" target=\"www.orau.gov\/2021ascr-cybersecurity\">ASCR Workshop on Cybersecurity and Privacy for Scientific  Computing Ecosystems<\/a><\/i>, November 3-5, 2021. <a href=\"javascript:showAbstract('Modern critical infrastructures (CI) and scientific computing ecosystems (SCE) are complex and vulnerable. The complexity of CI\/SCE, such as the distributed workload found across ASCR scientific computing facilities, does not allow for easy differentiation between emerging cyber security and reliability threats. It is also not easy to correctly identify the misbehaving systems. Sometimes, system failures are just caused by unintentional user misbehavior or actual hardware\/software reliability issues, but it may take some significant amount of time and effort to develop that understanding through root-cause analysis. On the security front, CI\/SCE are vital assets. They are prime targets of, and are vulnerable to, malicious cyber-attacks. Within DoE, inter-disciplinary and cross-facility collaboration (e.g., ORNL INTERSECT initiative, next-gen supercomputing OLCF6), traditional perimeter-based defense and demarcation line between malicious cyber-attacks and non-malicious system faults are blurring. Amidst realistic reliability and security threats, the ability to effectively distinguish between non-malicious faults and malicious attacks is critical not only in root cause identification but also in countermeasures generation. ');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/li21toward.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#li21toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Hal Finkel, Pete Beckman, Christian Engelmann, Shantenu Jha, and Jack Lange. <b>Research Opportunities in Operating Systems for Scientific Edge Computing<\/b>. <i>White paper by the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/OSRoundtable2021\" target=\"www.orau.gov\/OSRoundtable2021\">ASCR Roundtable Discussions on Operating-Systems Research 2021<\/a><\/i>, January 25, 2021. <a href=\"javascript:showAbstract('As scientific experiments generate ever-increasing amounts of data, and grow in operational complexity, modern experimental science demands unprecedented computational capabilities at the edge - physically proximate to each experiment. While some requirements on these computational capabilities are shared with high-performance-computing (HPC) systems, scientific edge computing has a number of unique challenges. In the following, we survey current trends in system software and edge systems for scientific computing, associated research challenges and open questions, infrastructure requirements for operating-systems research, communities who should be involved in that research, and the anticipated benefits of success.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/finkel21research2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#finkel21research2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Hal Finkel, Pete Beckman, Ron Brightwell, Rudi Eigenmann, Christian Engelmann, Roberto Gioiosa, Kamil Iskra, Shantenu Jha, Jack Lange, Tapasya Patki, and Kevin Pedretti. <b>Research Opportunities in Operating Systems for High-Performance Scientific Computing<\/b>. <i>White paper by the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/OSRoundtable2021\" target=\"www.orau.gov\/OSRoundtable2021\">ASCR Roundtable Discussions on Operating-Systems Research 2021<\/a><\/i>, January 25, 2021. <a href=\"javascript:showAbstract('As high-performance-computing (HPC) systems continue to evolve, with increasingly diverse and heterogeneous hardware, increasingly-complex requirements for security and multi-tenancy, and increasingly-demanding requirements for resiliency and monitoring, research in operating systems must continue to seed innovation to meet future needs. In the following, we survey current trends in system software and HPC systems for scientific computing, associated research challenges and open questions, infrastructure requirements for operating-systems research, communities who should be involved in that research, and the anticipated benefits of success.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/finkel21research.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#finkel21research\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience by Codesign (and not as an Afterthought)<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/web.cvent.com\/event\/f64a4f28-b473-4808-924c-c8c3d9a2af63\/\" target=\"web.cvent.com\/event\/f64a4f28-b473-4808-924c-c8c3d9a2af63\/\">Workshop on Reimagining Codesign 2021<\/a><\/i>, March 16-18, 2021. <a href=\"javascript:showAbstract('Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale high-performance computing (HPC). Extreme heterogeneity, i.e., using multiple, and potentially configurable, types of processors, accelerators and memory\/storage in a single computing platform, will add a significant amount of complexity to the HPC hardware\/software eco-system. Hardware\/software HPC codesign for resilience is mostly nonexistent at this point! Resilience needs to become an integral part of the HPC hardware\/software ecosystem through codesign, such that the burden for resilience is on the system by design and not on the operator or user as an afterthought. Simply put, if resilience by design is not done now, in the early stages of extreme heterogeneity, the current state of practice for HPC resilience, global application-level checkpoint\/restart, will re-main the same for decades to come due to the high costs of adoption of alternatives later on. ');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21resilience2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann21resilience2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21resilience2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Petar Radojkovic, Manolis Marazakis, Paul Carpenter, Reiley Jeyapaul, Dimitris Gizopoulos, Martin Schulz, Adria Armejach, Eduard Ayguade, Fran&ccedil;ois Bodin, Ramon Canal, Franck Cappello, Fabien Chaix, Guillaume Colin de Verdiere, Said Derradji, Stefano Di Carlo, Christian Engelmann, Ignacio Laguna, Miquel Moreto, Onur Mutlu, Lazaros Papadopoulos, Olly Perks, Manolis Ploumidis, Bezhad Salami, Yanos Sazeides, Dimitrios Soudris, Yiannis Sourdis, Per Stenstrom, Samuel Thibault, Will Toms, and Osman Unsal. <b>Towards Resilient EU HPC Systems: A Blueprint<\/b>. <i>White paper by the <a href=\"http:\/\/resilienthpc.eu\" target=\"resilienthpc.eu\">European HPC resilience initiative<\/a><\/i>, April 9, 2020. <a href=\"javascript:showAbstract('This document aims to spearhead a Europe-wide discussion on HPC system resilience and to help the European HPC community define best practices for resilience. We analyse a wide range of state-of-the-art resilience mechanisms and recommend the most effective approaches to employ in large-scale HPC systems. Our guidelines will be useful in the allocation of available resources, as well as guiding researchers and research funding towards the enhancement of resilience approaches with the highest priority and utility. Although our work is focussed on the needs of next generation HPC systems in Europe, the principles and evaluations are applicable globally. This document is the first output of the ongoing European HPC resilience initiative and it covers individual nodes in HPC systems, encompassing CPU, memory, intra-node interconnect and emerging FPGA-based hardware accelerators. With community support and feedback on this initial document, we will update the analysis and expand the scope to include other types of accelerators, as well as networks and storage.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/radojkovic20towards.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#radojkovic20towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Rizwan Ashraf, and Saurabh Hukerikar. <b>Extreme Heterogeneity with Resilience by Design (and not as an Afterthought)<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/orau.gov\/exheterogeneity2018\/\" target=\"orau.gov\/exheterogeneity2018\/\">Extreme Heterogeneity Virtual Workshop 2018<\/a><\/i>, January 23-24, 2018. <a href=\"javascript:showAbstract('Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale high-performance computing (HPC). Extreme heterogeneity, i.e., using multiple, and potentially configurable, types of processors, accelerators and  memory\/storage in a single computing platform, will add a significant amount of complexity to the HPC hardware\/software ecosystem. The notion of correct computation and program state assumed by users and application developers today, which has been based on binary bit-level correctness, will no longer hold for processing elements based on quantum qubits and analog circuits that model spiking neurons in neuromorphic computing elements. The diverse set of compute and memory components in future heterogeneous systems will require novel hardware and software resilience solutions. Errors and failures reported by such heterogeneous hardware will need to be handled by the appropriate software component to enable efficient masking, recovery, and avoidance with little burden on the user. Similarly, errors and failures reported by the software running on such heterogeneous hardware need to be equally efficiently handled with little burden on the user. This requires a new approach, where resilience is holistically provided by the HPC hardware\/software ecosystem. The key challenges are to design and to operate extreme heterogeneous HPC systems with (1) wide-ranging resilience capabilities in system software, programming models, libraries, and applications, (2) interfaces and mechanisms for coordinating resilience capabilities across diverse hardware and software components, (3) appropriate metrics and tools for assessing performance, resilience, and energy, and (4) an understanding of the performance, resilience and energy trade-off that eventually results in well-informed HPC system design choices and runtime decisions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18extreme.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann18extreme\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Devesh Tiwari, Saurabh Gupta, and Christian Engelmann. <b>Lightweight, Actionable Analytical Tools Based on Statistical Learning for Efficient System Operations<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2016\" target=\"hpc.pnl.gov\/modsim\/2016\">Workshop on Modeling &#038; Simulation of Systems &#038; Applications (ModSim) 2016<\/a><\/i>, August 10-12, 2016. <a href=\"javascript:showAbstract('Modeling and simulation community has always relied on accurate and meaningful system data and parameters to drive analytical models and simulators. HPC systems continuously generate huge amount system event related data (e.g., system log, resource consumption log, RAS logs, power consumption logs), but meaningful interpretation and accuracy verification of such data is quite challenging. This talk offers a unique perspective and experience in demonstrating how modeling and simulation based research can actually be translated into production systems. We will discuss the short-term opportunities for modeling and simulation community to increase the impact and effectiveness of our analytical tools, &amp;#34;dos and don&amp;#39;ts&amp;#34;, long-term challenges and opportunities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tiwari16lightweight.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tiwari16lightweight.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tiwari16lightweight\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A Hardware\/Software Performance\/Resilience\/Power Co-Design Tool for Extreme-scale Computing<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2013\" target=\"hpc.pnl.gov\/modsim\/2013\">Workshop on Modeling &#038; Simulation of Exascale Systems &#038; Applications (ModSim) 2013<\/a><\/i>, September 18-19, 2013. <a href=\"javascript:showAbstract('xSim is a simulation-based performance investigation toolkit that permits running high-performance computing (HPC) applications in a controlled environment with millions of concurrent execution threads, while observing application performance in a simulated extreme-scale system for hardware\/software co-design. The presented work details newly developed features for xSim that permit the injection of MPI process failures, the propagation\/detection\/notification of such failures within the simulation, and their handling using application-level checkpoint\/restart. The newly added features also offer user-level failure mitigation (ULFM) extensions at the simulated MPI layer to support algorithm-based fault tolerance (ABFT). The presented solution permits investigating performance under failure and failure handling of checkpoint\/restart and ABFT solutions. The newly enhanced xSim is the very first performance tool that supports these capabilities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13hardware.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann13hardware.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann13hardware\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Marc Snir, and Robert W. Wisniewski, Jacob A. Abraham, Sarita V. Adve, Saurabh Bagchi, Pavan Balaji, Bill Carlson, Andrew A. Chien, Pedro Diniz, Christian Engelmann, Rinku Gupta, Fred Johnson, Jim Belak, Pradip Bose, Franck Cappello, Paul Coteus, Nathan A. Debardeleben, Mattan Erez, Saverio Fazzari, Al Geist, Sriram Krishnamoorthy, Sven Leyffer, Dean Liberty, Subhasish Mitra, Todd Munson, Rob Schreiber, Jon Stearley, and Eric Van Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i>Workshop report<\/i>, August 4-11, 2013. <a href=\"publications\/snir13addressing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir13addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Al Geist, Bob Lucas, Marc Snir, Shekhar Borkar, Eric Roman, Mootaz Elnozahy, Bert Still, Andrew Chien, Robert Clay, John Wu, Christian Engelmann, Nathan DeBardeleben, Rob Ross, Larry Kaplan, Martin Schulz, Mike Heroux, Sriram Krishnamoorthy, Lucy Nowell, Abhinav Vishnu, and Lee-Ann Talley. <b>U.S. Department of Energy Fault Management Workshop<\/b>. <i>Workshop report for the U.S. Department of Energy<\/i>, June 6, 2012. <a href=\"javascript:showAbstract('A Department of Energy (DOE) Fault Management Workshop was held on June 6, 2012 at the BWI Airport Marriot hotel in Maryland. The goals of this workshop were to: 1. Describe the required HPC resilience for critical DOE mission needs; 2. Detail what HPC resilience research is already being done at the DOE national laboratories and is expected to be done by industry or other groups; 3. Determine what fault management research is a priority for DOE&amp;#39;s Office of Science and National Nuclear Security Administration (NNSA) over the next five years; 4. Develop a roadmap for getting the necessary research accomplished in the timeframe when it will be needed by the large computing facilities across DOE.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/geist12department.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#geist12department\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A Performance\/Resilience\/Power Co-design Tool for Extreme-scale High-Performance Computing<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2012\" target=\"hpc.pnl.gov\/modsim\/2012\">Workshop on Modeling &#038; Simulation of Exascale Systems &#038; Applications (ModSim) 2012<\/a><\/i>, August 9-10, 2012. <a href=\"javascript:showAbstract('Performance, resilience and power consumption are key HPC system design factors that are highly interde-pendent. To enable extreme-scale computing it is essential to perform HPC hardware\/software co-design that identifies the cost\/benefit trade-off between these design factors for potential future architecture choices. The proposed research and development aims at developing an HPC hardware\/software co-design toolkit for evaluating the resilience\/power\/performance cost\/benefit trade-off of future architecture choices. The approach focuses on extending a simulation-based performance investigation toolkit with advanced resilience and power modeling and simulation features, such as (i) fault injection mechanisms, (ii) fault propagation, isolation, and detection models, (i) fault avoidance, masking, and recovery simulation, and (iv) power consumption models.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12performance.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann12performance\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, and Frank Mueller. <b>Dynamic Self-Aware Runtime Software for Exascale Systems<\/b>. <i>White paper for the U.S. Department of Energy&#39;s <a href=\"http:\/\/collab.cels.anl.gov\/display\/exaosr\/Position+Papers\" target=\"collab.cels.anl.gov\/display\/exaosr\/Position+Papers\">Exascale Operating Systems and Runtime Technical Council<\/a><\/i>, July 1, 2012. <a href=\"javascript:showAbstract('At exascale, the power consumption, resilience, and load balancing constraints, especially their dynamic nature and interdependence, and the scale of the system require a radical change in future high-performance computing (HPC) operating systems and runtimes (OS\/Rs). In contrast to the existing static OS\/R solutions, an exascale OS\/R is needed that is aware of the dynamically changing resources, constraints, and application needs, and that is able to autonomously coordinate (sometimes conflicting) responses to different changes in the system, simultaneously and at scale. To provide awareness and autonomic management, a novel, scalable and self-aware OS\/R is needed that becomes the brains of the entire X-stack. It dynamically analyzes past, current, and future system status and application needs. It optimizes system usage by scheduling, migrating, and restarting tasks within and across nodes as needed to deal with multi-dimensional constraints, such as power consumption, permanent and transient faults, resource degradation, heterogeneity, data locality, and load balance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12dynamic.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann12dynamic.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann12dynamic\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, and David E. Bernholdt. <b>Unified Execution Environment<\/b>. <i>White paper for the U.S. Department of Energy&#39;s <a href=\"http:\/\/collab.cels.anl.gov\/display\/exaosr\/Position+Papers\" target=\"collab.cels.anl.gov\/display\/exaosr\/Position+Papers\">Exascale Operating Systems and Runtime Technical Council<\/a><\/i>, July 1, 2012. <a href=\"publications\/vallee12unified.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#vallee12unified\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Nathan DeBardeleben, James Laros, John T. Daly, Stephen L. Scott, Christian Engelmann, and Bill Harrod. <b>High-End Computing Resilience: Analysis of Issues Facing the HEC Community and Path-Forward for Research and Development<\/b>. <i>White paper for the U.S. National Science Foundation&#39;s High-end Computing Program<\/i>, December 1, 2009. <a href=\"publications\/debardeleben09high-end.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#debardeleben09high-end\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"reports\"><\/a>Technical Reports<\/h4>\n<ol>\n<li>Brian Etz, Oral, Sarp, Rafael Ferreira Da Silva, Ryan Adamson, Anees Alnajjar, Tom Beck, Ashley Barker, Michael Brim, Paul Bryant, Christian Engelmann, Anjus George, Samuel Herts, Gustav Jansen, Rajesh Kalyanam, Ahmad Maroof Karimi, Jack Lange, Kellen Leland, Ketan Maheshwari, Marshall McDonnell, Bronson Messer II, Ross Miller, Daniel S. Pelfrey, Suzanne Prentice, Bran Radovanovic, David Rogers, Daniel Rosendo, A.J. Ruckman, Mallikarjun (Arjun) Shankar, Amir Shehata, Tyler Skluzacek, Renan Santos Souza, Veronica Melesse Vergar, Feiyi Wang, Jordan Webb, Patrick Widener, and Christopher Zimmer. <b>OLCF&#39;s Advanced Computing Ecosystem (ACE): FY25 Update for Ongoing Efforts<\/b>. Technical Report, ORNL\/TM-2025\/4050, Oak Ridge National Laboratory, November 30, 2025. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/3006499\" target=\"publication\">10.2172\/3006499<\/a>. <a href=\"publications\/etz25olcf.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#etz25olcf\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rafael Ferreira da Silva, Robert Moore, Benjamin Mintz, Rigoberto Advincula, Anees Alnajjar, Luke Baldwin, Craig Bridges, Ryan Coffee, Ewa Deelman, Christian Engelmann, Brian Etz, Millie Firestone, Ian Foster, Panchapakesan Ganesh, Leslie Hamilton, Dale Huber, Ilia Ivanov, Shantenu Jha, Ying Li, Yongtao Liu, Jay Lofstead, Anirban Mandal, Hector Martin, Theresa Mayer, Marshall McDonnell, Vijayakumar Murugesan, Sal Nimer, Nageswara Rao, Martin Seifrid, Mitra Taheri, Michela Taufer, and Konstantinos Vogiatzis. <b>Shaping the Future of Self-Driving Autonomous Laboratories Workshop<\/b>. Technical Report, ORNL\/TM-2024\/3714, Oak Ridge National Laboratory, January 2, 2024. DOI <a href=\"http:\/\/dx.doi.org\/10.5281\/zenodo.14430232\" target=\"publication\">10.5281\/zenodo.14430232<\/a>. <a href=\"javascript:showAbstract('The Shaping the Future of Self-Driving Autonomous Laboratories workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/dasilva24shaping.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#dasilva24shaping\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Michael Brim and Christian Engelmann. <b>INTERSECT Architecture Specification: Microservice Architecture (Version 0.9)<\/b>. Technical Report, ORNL\/TM-2023\/3171, Oak Ridge National Laboratory, Oak Ridge, TN, USA, September 30, 2023. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/2333815\" target=\"publication\">10.2172\/2333815<\/a>. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL)&amp;#39;s Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) architecture project, titled &amp;#34;An Open Federated Architecture for the Laboratory of the Future&amp;#34;, creates an open federated hardware\/software architecture for the laboratory of the future using a novel system of systems (SoS) and microservice architecture approach, connecting scientific instruments, robot-controlled laboratories and edge\/center computing\/data resources to enable autonomous experiments, ``self-driving'' laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The project describes science use cases as design patterns that identify and abstract the involved hardware\/software components and their interactions in terms of control, work and data flow. It creates a SoS architecture of the federated hardware\/software ecosystem that clarifies terms, architectural elements, the interactions between them and compliance. It further designs a federated microservice architecture, mapping science use case design patterns to the SoS architecture with loosely coupled microservices, standardized interfaces and multi programming language support. The primary deliverable of this project is an INTERSECT Open Architecture Specification, containing the science use case design pattern catalog, the federated SoS architecture specification and the federated microservice architecture specification. This document represents the microservice architecture of the INTERSECT Open Architecture Specification.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/brim23microservice.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#brim23microservice\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Suhas Somnath. <b>INTERSECT Architecture Specification: Use Case Design Patterns (Version 0.9)<\/b>. Technical Report, ORNL\/TM-2023\/3133, Oak Ridge National Laboratory, Oak Ridge, TN, USA, September 30, 2023. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/2229218\" target=\"publication\">10.2172\/2229218<\/a>. <a href=\"javascript:showAbstract('Connecting scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing (HPC) center enables autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) Open Architecture enables science breakthroughs using intelligent networked systems, instruments and facilities with a federated hardware\/software architecture for the laboratory of the future. It relies on a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture. This document introduces the science use case design patterns of the INTERSECT Architecture. It describes the overall background, the involved terminology and concepts, and the pattern format and classification. It further details the 12 defined patterns and provides insight into building solutions from these patterns. The document also describes the application of these patterns in the context of several INTERSECT autonomous laboratories. The target audience are computer, computational, instrument and domain science experts working in the field of autonomous experiments.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23use.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann23use\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Rizwan Ashraf, Saurabh Hukerikar, Mohit Kumar, and Piyush Sao. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (Version 2.0)<\/b>. Technical Report, ORNL\/TM-2022\/2809, Oak Ridge National Laboratory, Oak Ridge, TN, USA, December 16, 2022. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/1922296\" target=\"publication\">10.2172\/1922296<\/a>. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance &amp; power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers&amp;#39; understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In this version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22rdp-20.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann22rdp-20\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Michael Brim and Christian Engelmann. <b>INTERSECT Architecture Specification: Microservice Architecture (Version 0.5)<\/b>. Technical Report, ORNL\/TM-2022\/2715, Oak Ridge National Laboratory, Oak Ridge, TN, USA, September 30, 2022. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/1902805\" target=\"publication\">10.2172\/1902805<\/a>. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL)&amp;#39;s Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) architecture project, titled &amp;#34;An Open Federated Architecture for the Laboratory of the Future&amp;#34;, creates an open federated hardware\/software architecture for the laboratory of the future using a novel system of systems (SoS) and microservice architecture approach, connecting scientific instruments, robot-controlled laboratories and edge\/center computing\/data resources to enable autonomous experiments, ``self-driving'' laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The project describes science use cases as design patterns that identify and abstract the involved hardware\/software components and their interactions in terms of control, work and data flow. It creates a SoS architecture of the federated hardware\/software ecosystem that clarifies terms, architectural elements, the interactions between them and compliance. It further designs a federated microservice architecture, mapping science use case design patterns to the SoS architecture with loosely coupled microservices, standardized interfaces and multi programming language support. The primary deliverable of this project is an INTERSECT Open Architecture Specification, containing the science use case design pattern catalog, the federated SoS architecture specification and the federated microservice architecture specification. This document represents the microservice architecture of the INTERSECT Open Architecture Specification.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/brim22microservice.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#brim22microservice\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Suhas Somnath. <b>INTERSECT Architecture Specification: Use Case Design Patterns (Version 0.5)<\/b>. Technical Report, ORNL\/TM-2022\/2681, Oak Ridge National Laboratory, Oak Ridge, TN, USA, September 30, 2022. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/1896984\" target=\"publication\">10.2172\/1896984<\/a>. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL)&amp;#39;s Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) architecture project, titled &amp;#34;An Open Federated Architecture for the Laboratory of the Future&amp;#34;, creates an open federated hardware\/software architecture for the laboratory of the future using a novel system of systems (SoS) and microservice architecture approach, connecting scientific instruments, robot-controlled laboratories and edge\/center computing\/data resources to enable autonomous experiments, ``self-driving'' laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The project describes science use cases as design patterns that identify and abstract the involved hardware\/software components and their interactions in terms of control, work and data flow. It creates a SoS architecture of the federated hardware\/software ecosystem that clarifies terms, architectural elements, the interactions between them and compliance. It further designs a federated microservice architecture, mapping science use case design patterns to the SoS architecture with loosely coupled microservices, standardized interfaces and multi programming language support. The primary deliverable of this project is an INTERSECT Open Architecture Specification, containing the science use case design pattern catalog, the federated SoS architecture specification and the federated microservice architecture specification. This document represents the science use case design pattern catalog of the INTERSECT Open Architecture Specification.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22use.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann22use\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (Version 1.2)<\/b>. Technical Report, ORNL\/TM-2017\/745, Oak Ridge National Laboratory, Oak Ridge, TN, USA, August 1, 2017. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/1436045\" target=\"publication\">10.2172\/1436045<\/a>. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance &amp; power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers&amp;#39; understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In this version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17rdp-12.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar17rdp-12\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (Version 1.1)<\/b>. Technical Report, ORNL\/TM-2016\/767, Oak Ridge National Laboratory, Oak Ridge, TN, USA, December 1, 2016. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/1345793\" target=\"publication\">10.2172\/1345793<\/a>. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore the resilience challenge for extreme-scale HPC systems requires management of various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in HPC systems future systems are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. As a result the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods and metrics to investigate and evaluate resilience holistically in HPC systems that consider impact scope, handling coverage, and performance &amp; power efficiency across the system stack. Additionally, few of the current approaches are portable to newer architectures and software environments that will be deployed on future systems. In this document, we develop a structured approach to the management of HPC resilience using the concept of resilience-based design patterns. A design pattern is a general repeatable solution to a commonly occurring problem. We identify the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. Each established solution is described in the form of a pattern that addresses concrete problems in the design of resilient systems. The complete catalog of resilience design patterns provides designers with reusable design elements. We also define a framework that enhances a designer&amp;#39;s understanding of the important constraints and opportunities for the design patterns to be implemented and deployed at various layers of the system stack. This design framework may be used to establish mechanisms and interfaces to coordinate flexible fault management across hardware and software components. The framework also supports optimization of the cost-benefit trade-offs among performance, resilience, and power consumption. The overall goal of this work is to enable a systematic methodology for the design and evaluation of resilience technologies in extreme-scale HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner in spite of frequent faults, errors, and failures of various types.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar16rdp-11.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar16rdp-11\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (Version 1.0)<\/b>. Technical Report, ORNL\/TM-2016\/687, Oak Ridge National Laboratory, Oak Ridge, TN, USA, October 1, 2016. DOI <a href=\"http:\/\/dx.doi.org\/10.2172\/1338552\" target=\"publication\">10.2172\/1338552<\/a>. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest that very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Practical limits on power consumption in HPC systems will require future systems to embrace innovative architectures, increasing the levels of hardware and software complexities. The resilience challenge for extreme-scale HPC systems requires management of various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. These techniques must seek to improve resilience at reasonable overheads to power consumption and performance. While the HPC community has developed various solutions, application-level as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods and metrics to investigate and evaluate resilience holistically in HPC systems that consider impact scope, handling coverage, and performance &amp; power efficiency across the system stack. Additionally, few of the current approaches are portable to newer architectures and software ecosystems, which are expected to be deployed on future systems. In this document, we develop a structured approach to the management of HPC resilience based on the concept of resilience-based design patterns. A design pattern is a general repeatable solution to a commonly occurring problem. We identify the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. The catalog of resilience design patterns provides designers with reusable design elements. We define a design framework that enhances our understanding of the important constraints and opportunities for solutions deployed at various layers of the system stack. The framework may be used to establish mechanisms and interfaces to coordinate flexible fault management across hardware and software components. The framework also enables optimization of the cost-benefit trade-offs among performance, resilience, and power consumption. The overall goal of this work is to enable a systematic methodology for the design and evaluation of resilience technologies in extreme-scale HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner in spite of frequent faults, errors, and failures of various types.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar16rdp-10.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar16rdp-10\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Frank Mueller, Christian Engelmann, Kurt Ferreira, Ron Brightwell, and Rolf Riesen. <b>Detection and Correction of Silent Data Corruption for Large-Scale High-Performance Computing<\/b>. Technical Report, ORNL\/TM-2012\/227, Oak Ridge National Laboratory, Oak Ridge, TN, USA, June 1, 2012. <a href=\"javascript:showAbstract('Faults have become the norm rather than the exception for high-end computing on clusters with 10s\/100s of thousands of cores. Exacerbating this situation, some of these faults remain undetected, manifesting themselves as silent errors that corrupt memory while applications continue to operate and report incorrect results. This paper studies the potential for redundancy to both detect and correct soft errors in MPI message-passing applications. Our study investigates the challenges inherent to detecting soft errors within MPI application while providing transparent MPI redundancy. By assuming a model wherein corruption in application data manifests itself by producing differing MPI message data between replicas, we study the best suited protocols for detecting and correcting MPI data that is the result of corruption. To experimentally validate our proposed detection and correction protocols, we introduce RedMPI, an MPI library which resides in the MPI profiling layer. RedMPI is capable of both online detection and correction of soft errors that occur in MPI applications without requiring any modifications to the application source by utilizing either double or triple redundancy. Our results indicate that our most efficient consistency protocol can successfully protect applications experiencing even high rates of silent data corruption with runtime overheads between 0% and 30% as compared to unprotected applications without redundancy. Using our fault injector within RedMPI, we observe that even a single soft error can have profound effects on running applications, causing a cascading pattern of corruption in most cases causes that spreads to all other processes. RedMPI&amp;#39;s protection has been shown to successfully mitigate the effects of soft errors while allowing applications to complete with correct results even in the face of errors.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala12detection.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#fiala12detection\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Hybrid Full\/Incremental Checkpoint\/Restart for MPI Jobs in HPC Environments<\/b>. Technical Report, ORNL\/TM-2010\/162, Oak Ridge National Laboratory, Oak Ridge, TN, USA, August 1, 2010. <a href=\"javascript:showAbstract('As the number of cores in high-performance computing environments keeps increasing, faults are becoming common place. Checkpointing addresses such faults but captures full process images even though only a subset of the process image changes between checkpoints. We have designed a high-performance hybrid disk-based full\/incremental checkpointing technique for MPI tasks to capture only data changed since the last checkpoint. Our implementation integrates new BLCR and LAM\/MPI features that complement traditional full checkpoints. This results in significantly reduced checkpoint sizes and overheads with only moderate increases in restart overhead. After accounting for cost and savings, benefits due to incremental checkpoints significantly outweigh the loss on restart operations. Experiments in a cluster with the NAS Parallel Benchmark suite and mpiBLAST indicate that savings due to replacing full checkpoints with incremental ones average 16.64 seconds while restore overhead amounts to just 1.17 seconds. These savings increase with the frequency of incremental checkpoints. Overall, our novel hybrid full\/incremental checkpointing is superior to prior non-hybrid techniques.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang10hybrid.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#wang10hybrid\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Proactive Process-Level Live Migration and Back Migration in HPC Environments<\/b>. Technical Report, ORNL\/TM-2010\/161, Oak Ridge National Laboratory, Oak Ridge, TN, USA, August 1, 2010. <a href=\"javascript:showAbstract('As the number of nodes in high-performance computing environments keeps increasing, faults are becoming common place. Reactive fault tolerance (FT) often does not scale due to massive I\/O requirements and relies on manual job resubmission. This work complements reactive with proactive FT at the process level. Through health monitoring, a subset of node failures can be anticipated when one&amp;#39;s health deteriorates. A novel process-level live migration mechanism suppor ts continued execution of applications during much of processes migration. This scheme is integrated into an MPI execution environment to transparently sustain health-inflicted node failures, which eradicates the need to restart and requeue MPI jobs. Experiments indicate that 1-6.5 seconds of prior warning are required to successfully trigger live process migration while similar operating system virtualization mechanisms require 13-24 seconds. This self-healing approach complements reactive FT by nearly cutting the number of checkpoints in half when 70% of the faults are handled proactively. The work also provides a novel back migration approach to eliminate load imbalance or bottlenecks caused by migrated tasks. Experiments indicate the larger the amount of outstanding execution, the higher the benefit due to back migration will be.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang10proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#wang10proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"reports\"><\/a>Datasets<\/h4>\n<ol>\n<li>Woong Shin, Vladyslav Oles, Anna Schmedding, George Ostrouchov, Evgenia Smirni, Christian Engelmann, and Feiyi Wang. <b>OLCF Summit Supercomputer GPU Snapshots During Double-Bit Errors and Normal Operations<\/b>. Dataset, April 20, 2023. DOI <a href=\"http:\/\/dx.doi.org\/10.13139\/OLCF\/1970187\" target=\"publication\">10.13139\/OLCF\/1970187<\/a>. <a href=\"javascript:showAbstract('As we move into the exascale era, the power and energy footprints of high-performance computing (HPC) systems have grown significantly larger. Due to the harsh power and thermal conditions the system, components are exposed to extreme operating conditions. Operation of such modern HPC systems requires deep insights into long term system behavior to maintain its efficiency as well as its longevity. To help the HPC community to gain such insights, we provide double-bit errors using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). The dataset relies on Nvidia XID records internally collected by GPU firmware at the time of failure occurrence, on the reboot-time logs of each Summit node, on node-level job scheduler records collected after each job termination, and on a 1Hz data rate from the baseboard management controllers (BMCs) of each Summit compute node using the OpenBMC event subscription protocol.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#shin23olcf\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mallikarjun Shankar, George Ostrouchov, Don Maxwell, James Rogers, Rizwan Ashraf, and Christian Engelmann. <b>GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability<\/b>. Dataset, September 2, 2020. DOI <a href=\"http:\/\/dx.doi.org\/10.13139\/ORNLNCCS\/1657202\" target=\"publication\">10.13139\/ORNLNCCS\/1657202<\/a>. <a href=\"javascript:showAbstract('George Ostrouchov, Don Maxwell, Rizwan Ashraf, Mallikarjun Shankar, and James Rogers. 2020. GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC &amp;#39;20). Association for Computing Machinery, New York, NY, USA. Data and code for SC20 paper about Titan GPU reliability analysis: https:\/\/github.com\/olcf\/TitanGPULife. Includes R  code to generate graphics for paper and additional analyses.  See code\/README for instructions. Includes original Titan GPU reliability data on over 100,000 collective hours of operation: data\/titan.gpu.history.txt - history data, data\/titan.service.txt - service nodes for exclusion. Includes output data files produced by code\/TitanGPUmodel.Rmd: data\/gc_full.csv - cleaned up data (see paper and R code); data\/gc_summary_loc.csv - one record per GPU (variables: SN, time, nlife, nloc, last, col, row, cage, slot, node, max_loc_events, time_max_loc, dbe, dbe_loc, otb, otb_loc, out, batch, days, years, dead, dead_otb, dead_dbe) (see paper and R code). Includes .Rmd analysis document as TitanGPUmode.html. Includes Python code to process data\/gc_full.csv into graphics from time-between-failure analyses: See code\/tbf-analyses\/README for instructions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#shankar20gpu\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"talks\"><\/a>Talks and Lectures<\/h4>\n<ol>\n<li>Christian Engelmann. <b>Towards a Strategy for Future Research Infrastructures<\/b>. Invited panelist at a Birds of a Feather session at the <a href=\"http:\/\/www.isc-hpc.com\" target=\"www.isc-hpc.com\">40th ISC High Performance Conference (ISC) 2025<\/a>, Hamburg, Germany, June 12, 2025. <a href=\"javascript:showAbstract('The amount of data gathered, shared and analysed in frontier research is set to increase dramatically in the coming decade, leading to unprecedented data processing, simulation\/prediction and analysis needs. As prime examples, the High Energy Physics and Radio Astronomy communities are gearing up to operate groundbreaking instruments such as the High-Luminosity Large Hadron Collider (LHC) and the Square Kilometer Array (SKA) , which will need data and compute capabilities many times larger than the currently available resources. Given the data volumes produced by these instruments, the size of the associated scientific communities and the scale of the analysis and computation problems, it is clear that distributed infrastructures integrating Edge, Cloud and large HPC\/AI centres into a data and compute continuum will be required.. This BoF will bring together top-level domain expert representatives from the High Energy Physics and Radio Astronomy domains and top-tier High Performance Computing infrastructure representatives across Europe and the US. Feedback from ISC community will be fed into the technical blueprint of the capabilities of the future infrastructure together with its roadmap for research, innovation and deployment of the future infrastructure');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann25towards.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann25towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Two Worlds Collide: Trustworthiness and Sustainability for Coupled HPC and AI Simulation<\/b>. Invited panelist at a Birds of a Feather session at the <a href=\"http:\/\/www.isc-hpc.com\" target=\"www.isc-hpc.com\">40th ISC High Performance Conference (ISC) 2025<\/a>, Hamburg, Germany, June 12, 2025. <a href=\"javascript:showAbstract('The &amp;#34;Two Worlds Collide Birds&amp;#34; of a Feather (BoF) series focuses on the experiences, challenges, and opportunities faced by laboratories and vendors in integrating deep learning (DL) and artificial intelligence (AI) with high-performance computing (HPC) for advanced simulation research. This fourth installment, titled ``Trustworthiness and Sustainability for Converged HPC and AI Simulation&amp;#39;' aims to promote a trustworthy and assured integration between established HPC simulation and the rapidly evolving DL ecosystem. Furthermore, this BoF seeks to address the emerging sustainability concerns associated with the verification and validation of converged HPC and AI simulations.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann25two.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann25two\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Federated Computing Environment for Autonomous Smart Laboratories<\/b>. Invited talk at the <a href=\"http:\/\/sos27.cscs.ch\" target=\"sos27.cscs.ch\">27th Workshop on Distributed Supercomputing (SOS) 2025<\/a>, Engelberg, Switzerland, March 20, 2025. <a href=\"javascript:showAbstract('The open Interconnected Science Ecosystem (INTERSECT) architecture connects scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing center to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery and evaluation. Its a novel approach consists of science use case design patterns, a system of systems architecture, and a microservice architecture. Failure resilience in federated ecosystems for instrument science is a critical challenge. Failures disrupt experiments and make them potentially useless, wasting valuable instrument, network and computing allocations and creating setbacks for scientists. A diverse, yet resilient, federated high-performance computing ecosystem is needed with traditional and accelerated capacity and capability computing resources and proper network and data storage resources, in part with on-demand and real-time features. This talk presents an overview of the resilient INTERSECT architecture, illustrates a resilient autonomous additive manufacturing use case, and discusses the future needs for incorporating such computational workloads into high-performance computing systems and facilities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann25federated.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann25federated\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Interconnected Science Ecosystem (INTERSECT)<\/b>. Invited talk at the <a href=\"http:\/\/www.hartree.stfc.ac.uk\" target=\"www.hartree.stfc.ac.uk\">Hartree Centre, Science and Technology Facilities Council, Daresbury, UK<\/a>, October 4, 2023. <a href=\"javascript:showAbstract('The Interconnected Science Ecosystem (INTERSECT) Initiative at Oak Ridge National Laboratory is in the process of creating an open federated hardware\/software architecture for the laboratory of the future, connecting scientific instruments, robot-controlled laboratories, and edge\/center computing\/data resources to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery, and evaluation. Its novel approach describes science use cases as design patterns that identify and abstract the involved hardware\/software components and their interactions in terms of control, work, and data flow. It creates a system-of-systems architecture of the federated hardware\/software ecosystem that clarifies terms, architectural elements, the interactions between them and compliance. It further designs a federated microservice architecture, mapping science use case design patterns to the system-of-systems architecture with loosely coupled microservices and standardized interfaces. The INTERSECT Open Architecture Specification contains a use case design pattern catalog, a federated system-of-systems architecture specification, and a federated microservice architecture specification. It is currently being used to prototype and deploy autonomous experiments and self-driving laboratories at Oak Ridge National Laboratory in the following science areas: (1) automation for electric grid interconnected-laboratory emulation\/simulation, (2) autonomous additive manufacturing, (3) autonomous continuous flow reactor synthesis, (4) autonomous electron microscopy, (5) autonomous robotic-controlled chemistry laboratory, and (6) integrating an ion trap quantum computing resource.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23interconnected4.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann23interconnected4\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Interconnected Science Ecosystem (INTERSECT) Architecture<\/b>. Invited talk at the <a href=\"http:\/\/smc2023.ornl.gov\" target=\"smc2023.ornl.gov\">20th Smoky Mountains Computational Sciences &#038; Engineering Conference (SMC)<\/a>, Knoxville, TN, USA, August 21-23, 2023. <a href=\"javascript:showAbstract('The Interconnected Science Ecosystem (INTERSECT) Initiative at Oak Ridge National Laboratory is in the process of creating an open federated hardware\/software architecture for the laboratory of the future, connecting scientific instruments, robot-controlled laboratories, and edge\/center computing\/data resources to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery, and evaluation. Its novel approach describes science use cases as design patterns that identify and abstract the involved hardware\/software components and their interactions in terms of control, work, and data flow. It creates a system-of-systems architecture of the federated hardware\/software ecosystem that clarifies terms, architectural elements, the interactions between them and compliance. It further designs a federated microservice architecture, mapping science use case design patterns to the system-of-systems architecture with loosely coupled microservices and standardized interfaces. The INTERSECT Open Architecture Specification contains a use case design pattern catalog, a federated system-of-systems architecture specification, and a federated microservice architecture specification. It is currently being used to prototype and deploy autonomous experiments and self-driving laboratories at Oak Ridge National Laboratory in the following science areas: (1) automation for electric grid interconnected-laboratory emulation\/simulation, (2) autonomous additive manufacturing, (3) autonomous continuous flow reactor synthesis, (4) autonomous electron microscopy, (5) autonomous robotic-controlled chemistry laboratory, and (6) integrating an ion trap quantum computing resource.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23interconnected3.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann23interconnected3\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Interconnected Science Ecosystem (INTERSECT) Architecture<\/b>. Seminar at the <a href=\"http:\/\/www.lrz-muenchen.de\" target=\"www.lrz-muenchen.de\">Leibniz Rechenzentrum (LRZ)<\/a>, Garching, Germany, July 10, 2023. <a href=\"javascript:showAbstract('The Interconnected Science Ecosystem (INTERSECT) Initiative at Oak Ridge National Laboratory is in the process of creating an open federated hardware\/software architecture for the laboratory of the future, connecting scientific instruments, robot-controlled laboratories, and edge\/center computing\/data resources to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery, and evaluation. Its novel approach describes science use cases as design patterns that identify and abstract the involved hardware\/software components and their interactions in terms of control, work, and data flow. It creates a system-of-systems architecture of the federated hardware\/software ecosystem that clarifies terms, architectural elements, the interactions between them and compliance. It further designs a federated microservice architecture, mapping science use case design patterns to the system-of-systems architecture with loosely coupled microservices and standardized interfaces. The INTERSECT Open Architecture Specification contains a use case design pattern catalog, a federated system-of-systems architecture specification, and a federated microservice architecture specification. It is currently being used to prototype and deploy autonomous experiments and self-driving laboratories at Oak Ridge National Laboratory in the following science areas: (1) automation for electric grid interconnected-laboratory emulation\/simulation, (2) autonomous additive manufacturing, (3) autonomous continuous flow reactor synthesis, (4) autonomous electron microscopy, (5) autonomous robotic-controlled chemistry laboratory, and (6) integrating an ion trap quantum computing resource.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23interconnected2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann23interconnected2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Interconnected Science Ecosystem (INTERSECT) Architecture<\/b>. Invited talk at the <a href=\"http:\/\/esailworkshop.ornl.gov\" target=\"esailworkshop.ornl.gov\">1st Ecosystems for Smart Autonomous Interconnected  Labs (E-SAIL) Workshop<\/a>, held in conjunction with the  <a href=\"http:\/\/www.isc-hpc.com\" target=\"www.isc-hpc.com\">38th ISC High  Performance (ISC) 2023<\/a>, Hamburg, Germany, May 25, 2023. <a href=\"javascript:showAbstract('The open Interconnected Science Ecosystem (INTERSECT) architecture connects scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing center to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery and evaluation. Its a novel approach consists of science use case design patterns, a system of systems architecture, and a microservice architecture.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23interconnected.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann23interconnected\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Designing Smart and Resilient Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/conferences\/cm\/conference\/pp22\" target=\"www.siam.org\/conferences\/cm\/conference\/pp22\">20th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2022<\/a>, Seattle, WA, USA, February 23-26, 2022. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and in extrapolating this knowledge to future-generation systems. It also describes the path forward in machine-in-the-loop operational intelligence for smart computing systems, leveraging operational data analytics in a loop control that maximizes productivity and minimizes costs through adaptive autonomous operation for resilience.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22designing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann22designing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Ben Mintz, Christian Engelmann, Elke Arenholz, and Ryan Coffee. <b>Enabling Self-Driven Experiments for Science through an Interconnected Science Ecosystem (INTERSECT)<\/b>. Panel at the <a href=\"http:\/\/smc2021.ornl.gov\" target=\"smc2021.ornl.gov\">17th Smoky  Mountains Computational Sciences &#038; Engineering Conference  (SMC)<\/a>, October 20, 2021. <a href=\"?page_id=55#mintz21enabling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Faults, Errors and Failures in Extreme-Scale Supercomputers<\/b>. Keynote talk at the  <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\">14th  Workshop on Resiliency in High Performance Computing  (Resilience) in Clusters, Clouds, and Grids<\/a>, held in  conjunction with the <a href=\"http:\/\/europar2014.dcc.fc.up.pt\" target=\"europar2014.dcc.fc.up.pt\">27th European Conference on Parallel and Distributed  Computing (Euro-Par) 2021<\/a>, Lisbon, Portugal, August 30, 2021. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of reliability experiences with some of the largest supercomputers in the world and recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in these systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21faults.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21faults\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Resilience Problem in Extreme Scale Computing: Experiences and the Path Forward<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/conferences\/cm\/conference\/cse21\" target=\"www.siam.org\/conferences\/cm\/conference\/cse21\">SIAM Conference on Computational Science and Engineering (CSE) 2021<\/a>, Fort Worth, TX, USA, March 1-5, 2021. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of reliability experiences with some of the largest supercomputers in the world and recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in these systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Smart and Resilient Extreme-Scale Systems<\/b>. Invited talk at the  <a href=\"http:\/\/www.hipeac.net\/2021\/spring-virtual\/#\/program\/sessions\/7854\/\" target=\"www.hipeac.net\/2021\/spring-virtual\/#\/program\/sessions\/7854\/\">Workshop on Resilience in High Performance Computing  (RESILIENTHPC)<\/a>, held in conjunction with the  <a href=\"http:\/\/www.hipeac.net\/2021\" target=\"www.hipeac.net\/2021\">European Network on High-performance Embedded Architecture   and Compilation (HiPEAC) Conference 2021<\/a>, Budapest, Hungary, January 19, 2021. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and in extrapolating this knowledge to future-generation systems. It also describes the path forward in machine-in-the-loop operational intelligence for smart computing systems, leveraging operational data analytics in a loop control that maximizes productivity and minimizes costs through adaptive autonomous operation for resilience.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21smart.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21smart\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Resilience Problem in Extreme Scale Computing<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/conferences\/cm\/conference\/pp20\" target=\"www.siam.org\/conferences\/cm\/conference\/pp20\">19th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2020<\/a>, Seattle, WA, USA, February 12-15, 2020. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of recent achievements in developing a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and in extrapolating this knowledge to future-generation systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann20resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann20resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience in Parallel Programming Environments<\/b>. Invited talk at the <a href=\"http:\/\/iadac.github.io\/events\/adac8\" target=\"iadac.github.io\/events\/adac8\">8th Accelerated Data Analytics and Computing (ADAC) Institute Workshop<\/a>, Tokyo, Japan, October 30-31, 2019. <a href=\"javascript:showAbstract('Recent reliability issues with one of the fastest supercomputers in the world, Titan at Oak Ridge National Laboratory, demonstrated the need for resilience in large-scale heterogeneous computing. OpenMP currently does not address error and failure behavior. The presented work takes a first step toward resilience for heterogeneous systems by providing the concepts for resilient OpenMP offload to devices. Using real-world error and failure observations, this work describes the concepts and terminology for resilient OpenMP target offload, including error and failure classes and resilience strategies. It details the experienced general-purpose computing on graphics processing units errors and failures in Titan. It further proposes improvements in OpenMP, including a preliminary prototype design, to support resilient offload to devices for efficient handling of errors and failures in heterogeneous high-performance computing systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19resilience3.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19resilience3\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience by Design (and not as an Afterthought)<\/b>. Invited talk at the <a href=\"http:\/\/sos23.ornl.gov\/\" target=\"sos23.ornl.gov\/\">23rd Workshop on Distributed Supercomputing (SOS) 2019<\/a>, Asheville, NC, USA, March 26-29, 2018. <a href=\"javascript:showAbstract('Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale high-performance computing (HPC). The challenge is to build a reliable HPC system within a given cost budget that achieves the expected performance. Every generation of supercomputers deployed at Oak Ridge National Laboratory (ORNL) had to deal with expected and unexpected faults, errors and failures. While these supercomputers are designed to deal with expected issues, unexpected reliability problems can lead to severe degradation in operational capabilities. For example, ORNL&amp;#39;s Titan supercomputer experienced an unexpected increase in general-purpose graphics processing unit (GPGPU) failures between 2015 and 2017. At the peak of the problem, Titan was losing an average of 12 GPGPUs (and corresponding compute nodes) per day. Over 50% of its 18,688 GPGPUs had to be replaced. The system and the applications using it were never designed to handle such a high failure rate in an efficient manner. Other past unexpected reliability issues with supercomputers at US Department of Energy HPC centers were caused by early wear-out, dirty power, bad solder, other manufacturing issues, design errors in hardware, design errors in software and user errors. With the expected decrease in reliability due to component count increases, process technology challenges, hardware heterogeneity and software complexity, risk mitigation against unexpected issues is becoming paramount to ensure the success of future extreme-scale HPC systems. Resilience needs to be holistically provided by the HPC hardware\/software ecosystem. The key challenges are to design and to operate extreme HPC systems with (1) wide-ranging resilience capabilities in hardware, system software, programming models, libraries, and applications, (2) interfaces and mechanisms for coordinating resilience capabilities across diverse hardware and software components, (3) appropriate metrics and tools for assessing performance, resilience, and energy, and (4) an understanding of the performance, resilience and energy trade-off that eventually results in well-informed HPC system design choices and runtime decisions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19resilience2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19resilience2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience for Extreme Scale Systems: Understanding the Problem<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/cse19\/\" target=\"www.siam.org\/meetings\/cse19\/\">SIAM Conference on Computational Science and Engineering (CSE) 2019<\/a>, Spokane, WA, USA, February 25 &#8211; March 1, 2018. <a href=\"javascript:showAbstract('Resilience is one of the critical challenges of extreme-scale high-performance computing (HPC) systems, as component counts increase, individual component reliability decreases, and software complexity increases. Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. This talk provides an overview of the Catalog project, which develops a taxonomy, catalog and models that capture the observed and inferred fault, error, and failure conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, this project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Rizwan Ashraf. <b>Modeling and Simulation of Extreme-Scale Systems for Resilience by Design<\/b>. Invited talk at the <a href=\"http:\/\/www.bnl.gov\/modsim2018\" target=\"www.bnl.gov\/modsim2018\">Workshop on Modeling and Simulation of Systems and Applications<\/a>, Seattle, WA, USA, August 15-17, 2018. <a href=\"javascript:showAbstract('Resilience is a serious concern for extreme-scale high-performance computing (HPC). While the HPC community has developed various resilience solutions, the solution space remains fragmented. We created a structured approach to the design, evaluation and optimization of HPC resilience using the concept of design patterns. A design pattern describes a generalized solution to a repeatedly occurring problem. We identified the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. Each well-known solution that addresses a specific resilience challenge is described in the form of a design pattern. We developed a resilience design pattern specification, language and catalog, which can be used by system architects, system software and library developers, application programmers, as well as users and operators as essential building blocks when designing and deploying resilience solutions. The resilience design pattern approach provides a unique opportunity for design space exploration. As each resilience solution is abstracted as a pattern and each solution&amp;#39;s properties are defined by pattern parameters, vertical and horizontal pattern compositions can describe the resilience capabilities of an entire HPC system. This permits the investigation of beneficial or counterproductive interactions between patterns and of the performance, resilience, and power consumption trade-off between different pattern parameters and compositions. The ultimate goal is to make resilience an integral part of the HPC hardware\/software ecosystem by coordinating the various existing resilience solutions in a design space exploration process, such that the burden for providing resilience is on the system by design and not on the user as an afterthought. We are in the early stages of developing a novel design space exploration tool that enables this investigation using modeling and simulation. We developed performance and resilience models for each resilience design pattern. We also leverage results from the Catalog project, a collaborative effort between Oak Ridge National Laboratory, Argonne National Laboratory and Lawrence Livermore National Laboratory that developed models of the faults, errors and failures in today's HPC systems. We also leverage recent results from the same project by Lawrence Livermore National Laboratory in application reliability patterns. The planned research extends and combines this work to model the performance, resilience, and power consumption of an entire HPC system, initially at node-level granularity, and to simulate the dynamic interactions between deployed resilience solutions and the rest of the system. In the next iteration, finer-grain modeling and simulation, such as at the computational unit level, is used to increase accuracy. This work leverages the experience of the investigators in parallel discrete event simulation of extreme-scale systems, such as the Extreme-scale Simulator (xSim). The current state of the art in resilience modeling and simulation is fragmented as well. There is currently no such design space exploration tool. Instead, each resilience solution is typically investigated separately. There is only a small amount of work on multi-resilience solutions, including by the investigators. While there is work in investigating the performance\/resilience trade-off space, there is almost no work in including power consumption.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18modeling.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18modeling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Characterizing Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/pasc18.pasc-conference.org\" target=\"pasc18.pasc-conference.org\">Platform for Advanced Scientific Computing (PASC) Conference 2018<\/a>, Basel, Switzerland, July 2-4, 2018. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18characterizing2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18characterizing2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Characterizing Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/iadac.github.io\/adac6\" target=\"iadac.github.io\/adac6\">6th Accelerated Data Analytics and Computing (ADAC) Institute Workshop<\/a>, Zurich, Switzerland, June 20-21, 2018. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18characterizing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18characterizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Pattern-based Modeling of Fail-stop and Soft-error Resilience for Iterative Linear Solvers<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/pp18\/\" target=\"www.siam.org\/meetings\/pp18\/\">18th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2018<\/a>, Tokyo, Japan, March 7-10, 2018. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC). While the HPC community has developed various resilience solutions, the solution space remains fragmented. With this work, we develop a structured approach to the design, evaluation and optimization of HPC resilience using the concept of design patterns. We identify the problems caused by faults, errors and failures in HPC systems and the techniques used to deal with these events. Each well-known solution that addresses a specific resilience challenge is described in the form of a pattern. We develop a catalog of such resilience design patterns, which may be used by system architects, system software and tools developers, application programmers, as well as users and operators as essential building blocks when designing and deploying resilience solutions. We also develop a design framework that enhances a designer&amp;#39;s understanding the opportunities for integrating multiple patterns across layers of the system stack and the important constraints during implementation of the individual patterns. It is also useful for designing mechanisms and interfaces to coordinate flexible fault management across hardware and software components. The resilience patterns and the design framework also enable exploration and evaluation of design alternatives and support optimization of the cost-benefit trade-offs among performance, protection coverage, and power consumption of resilience solutions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18pattern-based.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18pattern-based\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/pp18\/\" target=\"www.siam.org\/meetings\/pp18\/\">18th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2018<\/a>, Tokyo, Japan, March 7-10, 2018. <a href=\"javascript:showAbstract('The reliability of high-performance computing (HPC) platforms is among the most critical challenges as systems continue to increase component counts, while the individual component reliability decreases and software complexity increases. While most resilience solutions are designed to address a specific fault model, HPC applications must contend with extremely high rates of faults from various sources with different levels of severity. Therefore, resilience for extreme-scale HPC systems and their applications requires an integrated approach, which leverages detection, containment and mitigation capabilities from different layers of the HPC environment. With this work, we propose an approach based on design patterns to explore a multi-level resilience solution that addresses silent data corruptions and process failures. The structured approach enables evaluation of the key components of a multi-level resilience solution using pattern performance models and systematically integrating the patterns into a complete solution by assessing the interplay between the patterns. We describe the design steps to develop a multi-level resilience solution for an iterative linear solver application that combines algorithmic resilience features of the solver with the fault tolerance primitives provided by ULFM MPI. Our results demonstrate the viability of designing HPC applications capable of surviving simultaneous injection of hard and soft errors in a performance efficient manner.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann18resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>A Catalog of Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/an17\/\" target=\"www.siam.org\/meetings\/an17\/\">SIAM Annual Meeting (AM) 2017<\/a>, Pittsburgh, PA, USA, July 10-14, 2017. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann17catalog2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann17catalog2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Characterizing Faults, Errors and Failures in Extreme-Scale Computing Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.isc-hpc.com\" target=\"www.isc-hpc.com\">International Supercomputing Conference (ISC) 2017<\/a>, Frankfurt am Main, Germany, June 16-22, 2017. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann17characterizing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann17characterizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>A Catalog of Faults, Errors, and Failures in Extreme-Scale Systems<\/b>. Invited talk at the <a href=\"http:\/\/icl.cs.utk.edu\/workshops\/scheduling2017\/\" target=\"icl.cs.utk.edu\/workshops\/scheduling2017\/\">12th Scheduling for Large Scale Systems Workshop (SLSSW) 2017<\/a>, Knoxville, TN, USA, May 24-26, 2017. <a href=\"javascript:showAbstract('Building a reliable supercomputer that achieves the expected performance within a given cost budget and providing efficiency and correctness during operation in the presence of faults, errors, and failures requires a full understanding of the resilience problem. The Catalog project develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current supercomputers and extrapolates this knowledge to future-generation systems. To date, the Catalog project has analyzed billions of node hours of system logs from supercomputers at Oak Ridge National Laboratory and Argonne National Laboratory. This talk provides an overview of our findings and lessons learned.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann17catalog.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann17catalog\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>The Missing High-Performance Computing Fault Model<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/pp16\/\" target=\"www.siam.org\/meetings\/pp16\/\">17th SIAM Conference on Parallel Processing for Scientific Computing (PP) 2016<\/a>, Paris, France, April 12-15, 2016. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges. Resilience is one of the most important challenges. This talk will present recent work in developing the missing high-performance computing (HPC) fault model. This effort identifies, categorizes and models the fault, error and failure properties of today&amp;#39;s HPC systems. It develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current systems and extrapolates this knowledge to exascale HPC systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16missing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann16missing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience Challenges and Solutions for Extreme-Scale Supercomputing<\/b>. Invited talk at the <a href=\"http:\/\/www.usna.edu\" target=\"www.usna.edu\">United  States Naval Academy<\/a>, Annapolis, MD, USA, February 18, 2016. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2022) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2022). This talk provides an overview of recent and ongoing resilience research and development activities at Oak Ridge National Laboratory in advanced checkpoint storage architectures, process-level incremental checkpoint\/restart, proactive fault tolerance using prediction-triggered process or virtual machine migration, MPI process-level software redundancy, and soft-error injection tools to study the vulnerability of science applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16resilience2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann16resilience2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Toward A Fault Model And Resilience Design Patterns For Extreme Scale Systems<\/b>. Keynote talk at the  <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2015\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2015\">8th  Workshop on Resiliency in High Performance Computing  (Resilience) in Clusters, Clouds, and Grids<\/a>, held in  conjunction with the <a href=\"http:\/\/europar2014.dcc.fc.up.pt\" target=\"europar2014.dcc.fc.up.pt\">21st European Conference on Parallel and Distributed  Computing (Euro-Par) 2015<\/a>, Vienna, Austria, August 24-28, 2015. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2022) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2022). This talk provides an overview of two recently funded projects. The Characterizing Faults, Errors, and Failures in Extreme-Scale Systems project identifies, categorizes and models the fault, error and failure properties of US Department of Energy high-performance computing (HPC) systems. It develops a fault taxonomy, catalog and models that capture the observed and inferred conditions in current systems and extrapolate this knowledge to exascale HPC systems. The Resilience Design Patterns project will increase the ability of scientific applications to reach accurate solutions in a timely and efficient manner. Using a novel design pattern concept, it identifies and evaluates repeatedly occurring resilience problems and coordinates solutions throughout high-performance computing hardware and software.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann15toward.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann15toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience Challenges and Solutions for Extreme-Scale Supercomputing<\/b>. Invited talk at the  19th Workshop on Distributed Supercomputing (SOS)   2015, Park City, UT, USA, March 2-5, 2015. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2022) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2022). This talk provides an overview of recent and ongoing resilience research and development activities at Oak Ridge National Laboratory in advanced checkpoint storage architectures, process-level incremental checkpoint\/restart, proactive fault tolerance using prediction-triggered process or virtual machine migration, MPI process-level software redundancy, and soft-error injection tools to study the vulnerability of science applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann15resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann15resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>xSim: The Extreme-scale Simulator<\/b>. Seminar at the <a href=\"http:\/\/www.lrz-muenchen.de\" target=\"www.lrz-muenchen.de\">Leibniz Rechenzentrum (LRZ)<\/a>, Garching, Germany, February 23, 2015. <a href=\"javascript:showAbstract('The path to exascale high-performance computing (HPC) poses several challenges related to power, performance, and resilience. Investigating the performance and resilience of parallel applications at scale on future architectures and the performance and resilience impact of different architecture choices is an important component of HPC hardware\/software co-design. Without having access to future architectures at scale, simulation provides an alternative. The Extreme-scale Simulator (xSim) is a performance investigation toolkit that permits running applications in a controlled environment with millions of concurrent execution threads, while observing performance and resilience in a simulated extreme-scale system. Using a lightweight parallel discrete event simulation, xSim executes a Message Passing Interface (MPI) application on a much smaller system in a highly oversubscribed fashion with a virtual wall clock time, such that performance data can be extracted based on a processor and a network model. xSim is designed like a traditional performance tool, as an interposition library that sits between the MPI application and the MPI library, using the MPI profiling interface. It has been run up to 134,217,728 (2^27) MPI ranks using a 960-core Linux cluster. xSim also permits the injection of MPI process failures, the propagation\/detection\/notification of such failures within the simulation, and their handling within the simulation using application-level checkpoint\/restart. Another feature provides user-level failure mitigation (ULFM) extensions at the simulated MPI layer to support algorithm-based fault tolerance (ABFT). xSim is the very first performance tool that supports ULFM and ABFT.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann15xsim.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann15xsim\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Supporting the Development of Resilient Message Passing Applications using Simulation<\/b>. Invited talk at the <a href=\"http:\/\/www.dagstuhl.de\/en\/program\/calendar\/semhp\/?semnr=14402\" target=\"www.dagstuhl.de\/en\/program\/calendar\/semhp\/?semnr=14402\">Dagstuhl Seminar on Resilience in Exascale Computing<\/a>, Schloss Dagstuhl, Wadern, Germany, September 28 &#8211; October 1, 2014. <a href=\"javascript:showAbstract('An emerging aspect of high-performance computing (HPC) hardware\/software co-design is investigating performance under failure. The presented work extends the Extreme-scale Simulator (xSim), which was designed for evaluating the performance of message passing interface (MPI) applications on future HPC architectures, with fault-tolerant MPI extensions proposed by the MPI Fault Tolerance Working Group. xSim permits running MPI applications with millions of concurrent MPI ranks, while observing application performance in a simulated extreme-scale system using a lightweight parallel discrete event simulation. The newly added features offer user-level failure mitigation (ULFM) extensions at the simulated MPI layer to support algorithm-based fault tolerance (ABFT). The presented solution permits investigating performance under failure and failure handling of ABFT solutions. The newly enhanced xSim is the very first performance tool that supports ULFM and ABFT.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann14supporting.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann14supporting\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience Challenges and Solutions for Extreme-Scale Supercomputing<\/b>. Invited talk at the Technical University of Dresden,  Dresden, Germany, September 3, 2013. <a href=\"javascript:showAbstract('With the recent deployment of the 18 PFlop\/s Titan supercomputer and the exascale roadmap targeting 100, 300, and eventually 1,000 PFlop\/s by 2022, Oak Ridge National Laboratory is at the forefront of scientific capability computing. The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2022) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2022). This talk provides an overview of recent and ongoing resilience research and development activities at Oak Ridge National Laboratory in advanced checkpoint storage architectures, process-level incremental checkpoint\/restart, proactive fault tolerance using prediction-triggered process or virtual machine migration, MPI process-level software redundancy, and soft-error injection tools to study the vulnerability of science applications and of CMOS logic in processors and memory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann13resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Fault Tolerance Session<\/b>. Invited talk at the  <a href=\"http:\/\/www.aanmelder.nl\/exachallenge\" target=\"www.aanmelder.nl\/exachallenge\">The ExaChallenge Symposium<\/a>, Dublin, Ireland, October 16-17, 2012. <a href=\"publications\/engelmann12fault.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann12fault\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High-End Computing Resilience: Analysis of Issues Facing the HEC Community and Path Forward for Research and Development<\/b>. Invited talk at the Argonne National Laboratory (ANL)  Institute of Computing in Science (ICiS)  <a href=\"http:\/\/www.icis.anl.gov\/programs\/summer2012-4b\" target=\"www.icis.anl.gov\/programs\/summer2012-4b\">Summer Workshop Week on Addressing Failures in Exascale   Computing<\/a>, Park City, UT, USA, August 4-11, 2012. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2020) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2020). To provide input for a discussion of future needs in resilience research, development, and standards work, this talk gives a brief summary of the outcomes from the National HPC Workshop on Resilience, held in Arlington, VA, USA on August 12-14, 2009.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12high-end.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann12high-end\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience for Permanent, Transient, and Undetected Errors<\/b>. Invited talk at the  <a href=\"http:\/\/www.cs.sandia.gov\/Conferences\/SOS16\" target=\"www.cs.sandia.gov\/Conferences\/SOS16\">16th Workshop on Distributed Supercomputing (SOS)   2012<\/a>, Santa Barbara, CA, USA, March 12-15, 2012. <a href=\"javascript:showAbstract('With the ongoing deployment of 10-20 PFlop\/s supercomputers and the exascale roadmap targeting 100, 300, and eventually 1,000 PFlop\/s by 2020, the path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count (100,000-1,000,000 nodes with 1,000-10,000 cores per node by 2020) and component reliability decreases (7 nm technology with near-threshold voltage operation by 2020). This talk provides an overview of recent and ongoing resilience research and development activities at Oak Ridge National Laboratory, and of future needs in resilience research, development, and standards work.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann12resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Scaling To A Million Cores And Beyond: A Basic Understanding Of The Challenges Ahead On The Road To Exascale<\/b>. Invited talk at the <a href=\"http:\/\/researcher.ibm.com\/researcher\/view_page.php?id=2580\" target=\"researcher.ibm.com\/researcher\/view_page.php?id=2580\">1st International Workshop on Extreme Scale Parallel Architectures and Systems (ESPAS) 2012<\/a>, in conjunction with the <a href=\"http:\/\/www.hipeac.net\/conference\/paris\" target=\"www.hipeac.net\/conference\/paris\">7th International Conference on High-Performance and Embedded Architectures and Compilers (HiPEAC) 2012<\/a>, Paris France, January 24, 2012. <a href=\"javascript:showAbstract('On the road toward multi-petascale and exascale HPC, the trend in architecture goes clearly in only one direction. HPC systems will dramatically scale up in compute node and processor core counts. By 2020, an exascale system may have up to 1,000,000 compute nodes with 1,000 cores per node. The substantial growth in concurrency causes parallel application scalability issues due to sequential application parts, synchronizing communication, and other bottlenecks. Investigating parallel algorithm performance properties at this scale and with these architectural properties for HPC hardware\/software co-design is crucial to enable extreme-scale computing. The presented work utilizes the Extreme-scale Simulator (xSim) performance investigation toolkit to identify the scaling characteristics of a simple Monte Carlo algorithm from 1 to 16 million MPI processes on different multi-core architecture choices. The results show the limitations of strong scaling and the negative impact of employing more but less powerful cores for energy savings.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12scaling.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann12scaling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilient Software for ExaScale Computing<\/b>. Invited talk at the Birds of a Feather Session on Resilient Software for ExaScale Computing at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">24th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2011<\/a>, Seattle, WA, USA, November 17, 2011. <a href=\"javascript:showAbstract('ExaScale computing systems will likely consist of millions of cores executing applications with billions of threads, based on 14nm or less CMOS technology, according to the ITRS roadmap. Processing elements built on this technology, coupled with dynamic power management will exhibit high variability in performance, between cores and across different runs. Even worse, preliminary figures indicates that on average about every couple of minutes - at least - something in the system will break. Traditional checkpointing strategies are unlikely to work, given the time it will take to save the huge quantities of data combined with the fact that they will need to be restored frequently. This BoF wants to investigate resilient software: software that is able to survive failing hardware and continue to run, without minimal performance impact. Furthermore, we may also discuss tradeoffs between rerunning the application and the cost of instrumentation to deal with resilience.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann11resilient.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann11resilient\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience and Hardware\/Software Co-design for Extreme-Scale Supercomputing<\/b>. Seminar at the <a href=\"http:\/\/www.bsc.es\" target=\"www.bsc.es\">Barcelona Supercomputing Center<\/a>, Barcelona, Spain, July 27, 2011. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL) provides the most powerful high-performance computing (HPC) resources in the world for open scientific research. Jaguar, a 224,162-core Cray XT5 with a LINPACK performance of 1.759 PFlop\/s, for example, is the world&amp;#39;s 3rd fastest supercomputer. 80% of its resources are allocated through a reviewed process to address the most challenging scientific problems in climate modeling, renewable energy, materials science, fusion and other areas. ORNL's Computer Science and Mathematics Division performs computer science and mathematics research to increase supercomputer efficiency and application scientist productivity while accelerating time to solution for scientific breakthroughs. This talk details recent research advancements at ORNL in two areas: (1) resilience and (2) hardware\/software co-design for extreme-scale supercomputing. Both are essential on the road toward exa-scale HPC systems with millions-to-billions of cores. Due to the expected drastic increase in scale, the corresponding decrease in system mean-time to interrupt warrants a rethinking of the traditional checkpoint\/restart approach for HPC resilience. New concepts discussed in this talk range from preventative measures, such as task migration based on fault prediction, to more aggressive fault masking, such as various levels of redundancy. Further, the expected drastic increase in task parallelism requires redesigning algorithms to avoid the consequences of Amdahl's law at extreme scale. As million-way task parallel systems don't exist yet, this talk discusses a lightweight system simulation approach for performance estimation of algorithms at scale.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann11resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann11resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Scalable HPC System Monitoring<\/b>. Invited talk at the 3rd HPC Resiliency Summit: Workshop on Resiliency for Petascale HPC 2010, in conjunction with the <a href=\"http:\/\/www.lanl.gov\/conferences\/lacss\/2010\" target=\"www.lanl.gov\/conferences\/lacss\/2010\">3rd Los Alamos Computer Science Symposium (LACSS) 2010<\/a>, Santa Fe, NM, USA, October 13, 2010. <a href=\"javascript:showAbstract('We present a monitoring system for large-scale parallel and distributed computing environments that allows to trade-off accuracy in a tunable fashion to gain scalability without compromising fidelity. The approach relies on classifying each gathered monitoring metric based on individual needs and on aggregating messages containing classes of individual monitoring metrics using a tree-based overlay network. The MRNet-based prototype is able to significantly reduce the amount of gathered and stored monitoring data, e.g., by a factor of  56 in comparison to the Ganglia distributed monitoring system. A simple scaling study reveals, however, that further efforts are needed in reducing the amount of data to monitor future-generation extreme-scale systems with up to 1,000,000 nodes. The implemented solution did not had a measurable performance impact as the 32-node test system did not produce enough monitoring data to interfere with running applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann10scalable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Beyond Application-Level Checkpoint\/Restart &#8211; Advanced Software Approaches for Fault Resilience<\/b>. Talk at the <a href=\"http:\/\/www.speedup.ch\/workshops\/w39_2010.html\" target=\"www.speedup.ch\/workshops\/w39_2010.html\">39th SPEEDUP Workshop on High Performance Computing<\/a>, Zurich, Switzerland, September 6, 2010. <a href=\"publications\/engelmann10beyond.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10beyond\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Stephen L. Scott. <b>Reliability, Availability, and Serviceability (RAS) for Petascale High-End Computing and Beyond<\/b>. Talk at the <a href=\"http:\/\/www.usenix.org\/events\/fastos10\" target=\"www.usenix.org\/events\/fastos10\">Forum to Address Scalable Technology for Runtime and Operating Systems (FAST-OS) Workshop<\/a>, in conjunction with the <a href=\"http:\/\/www.usenix.org\/events\/confweek10\" target=\"www.usenix.org\/events\/confweek10\">USENIX Federated Conferences Week (USENIX) 2010<\/a>, Boston MA, USA, June 22, 2010. <a href=\"javascript:showAbstract('This project aims at scalable technologies for providing high-level RAS for next-generation petascale scientific high-performance computing (HPC) resources and beyond as outlined by the U.S. Department of Energy (DOE) Forum to Address Scalable Technology for Runtime and Operating Systems (FAST-OS) and the U.S. National Coordination Office for Networking and Information Technology Research and Development (NCO\/NITRD) High-End Computing Revitalization Task Force (HECRTF) activities. Based on virtualized adaptation, reconfiguration, and preemptive measures, the ultimate goal is to provide for non-stop scientific computing on a 24x7 basis without interruption. The taken technical approach leverages system-level virtualization technology to enable transparent proactive and reactive fault tolerance mechanisms on extreme scale HPC systems. This effort targets: (1) reliability analysis for identifying pre-fault indicators, predicting failures, and modeling and monitoring component and system reliability, (2) proactive fault tolerance technology based on preemptive migration away from components that are about to fail, (3) reactive fault tolerance enhancements, such as checkpoint interval and placement adaptation to actual and predicted system health threats, and (4) holistic fault tolerance through combination of adaptive proactive and reactive fault tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann10reliability.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10reliability\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience Challenges at the Exascale<\/b>. Talk at the <a href=\"http:\/\/www.csm.ornl.gov\/workshops\/SOS14\" target=\"www.csm.ornl.gov\/workshops\/SOS14\">14th Workshop on Distributed Supercomputing (SOS) 2010<\/a>, Savannah, GA, USA, March 8-11, 2010. <a href=\"javascript:showAbstract('The path to exascale computing poses several research challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Resilience, i.e., providing efficiency and correctness in the presence of faults, is one of the most important exascale computer science challenges as systems scale up in component count and component reliability decreases. This talk discusses the future needs in resilience research, development, and standards work based on the outcomes from the National HPC Workshop on Resilience, held in Arlington, VA, USA on August 12-14, 2009.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann10resilience.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Stephen L. Scott. <b>HPC System Software Research at Oak Ridge National Laboratory<\/b>. Seminar at the <a href=\"http:\/\/www.lrz-muenchen.de\" target=\"www.lrz-muenchen.de\">Leibniz Rechenzentrum (LRZ)<\/a>, Garching, Germany, February 22, 2010. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL) is the largest energy laboratory in the United States. Its National Center for Computational Sciences (NCCS) provides the most powerful computing resources in the world for open scientific research. Jaguar, a Cray XT5 system at NCCS, is the fastest supercomputer in the world. It recently ranked #1 in the Top 500 List of Supercomputer Sites with a maximal LINPACK benchmark performance of 1.759 PFlop\/s and a theoretical peak performance of 2.331 PFlop\/s, where 1 PFlop\/s is 10^15 Floating Point Operations Per Second. Annually, 80 percent of Jaguar&amp;#39;s resources are allocated through the U.S Department of Energy's Innovative and Novel Computational Impact on Theory and Experiment (INCITE) program, a competitively selected, peer reviewed process open to researchers from universities, industry, government and non-profit organizations. These allocations address some of the most challenging scientific problems in areas such as climate modeling, renewable energy, materials science, fusion and combustion. In conjunction with NCCS, the Computer Science and Mathematics Division at ORNL performs basic and applied research in HPC, mathematics, and intelligent systems. This talk gives a summary of the HPC research and development in system software performed at ORNL, including resilience at extreme scale and virtualization technologies in HPC. Specifically, this talk will focus on advanced resilience technologies, such as migration of computation away from components that are about to fail and on management and customization of virtualized environments.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann10hpc.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10hpc\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High-Performance Computing Research Internship and Appointment Opportunities at Oak Ridge National Laboratory<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, December 14, 2009. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL) is the largest energy laboratory in the United States. Its National Center for Computational Sciences (NCCS) provides the most powerful computing resources in the world for open scientific research. Jaguar, a Cray XT5 system at NCCS, is the fastest supercomputer in the world. It recently ranked #1 in the Top 500 List of Supercomputer Sites with a maximal LINPACK benchmark performance of 1.759 PFlop\/s and a theoretical peak performance of 2.331 PFlop\/s, where 1 PFlop\/s is 10^15 Floating Point Operations Per Second. Annually, 80 percent of Jaguar&amp;#39;s resources are allocated through the U.S Department of Energy's Innovative and Novel Computational Impact on  Theory and Experiment (INCITE) program, a competitively selected, peer reviewed process open to researchers from universities, industry, government and non-profit organizations. These allocations address some of the most challenging scientific problems in areas such as climate modeling, renewable energy, materials science, fusion and combustion. In conjunction with NCCS, the Computer Science and Mathematics Division at ORNL performs basic and applied research in HPC, mathematics, and intelligent systems. This talk gives a summary of the HPC research performed at ORNL. It provides details about the Jaguar peta-scale computing resource, an overview of the computational science research carried out using ORNL's computing resources, and a description of various computer science efforts targeting solutions for next-generation HPC systems. This talk also provides information about internship opportunities for MSc students and research appointment opportunities for recent graduates.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09high2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09high2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>JCAS &#8211; IAA Simulation Efforts at Oak Ridge National Laboratory<\/b>. Invited talk at the <a href=\"http:\/\/www.cs.sandia.gov\/CSRI\/Workshops\/2009\/IAA\" target=\"www.cs.sandia.gov\/CSRI\/Workshops\/2009\/IAA\">IAA Workshop on HPC Architectural Simulation (HPCAS)<\/a>, Boulder, CO, USA, September 1-2, 2009. <a href=\"publications\/engelmann09jcas.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09jcas\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Modeling Techniques Towards Resilience<\/b>. Invited talk at the <a href=\"http:\/\/institute.lanl.gov\/resilience\/conferences\/2009\" target=\"institute.lanl.gov\/resilience\/conferences\/2009\">National HPC Workshop on Resilience 2009<\/a>, Arlington, VA, USA, August 12-14, 2009. <a href=\"publications\/engelmann09modeling.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09modeling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>System Resilience Research at ORNL in the Context of HPC<\/b>. Invited talk at the <a href=\"http:\/\/www.inria.fr\/inria\/organigramme\/fiche_ur-ren.fr.html\" target=\"www.inria.fr\/inria\/organigramme\/fiche_ur-ren.fr.html\">Institut National de Recherche en Informatique et en Automatique (INRIA)<\/a>, Rennes, France, May 15, 2009. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09system.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09system\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High-Performance Computing Research and MSc Internship Opportunities at Oak Ridge National Laboratory<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, May 11, 2009. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL) is the largest energy laboratory in the United States. Its National Center for Computational Sciences (NCCS) provides the most powerful computing resources in the world for open scientific research. Jaguar, a Cray XT5 system at NCCS, is the second HPC system to exceed 1 PFlop\/s (10^15 Floating Point Operations Per Second), and the fastest open science supercomputer in the world. It recently ranked #2 in the Top 500 List of Supercomputer Sites with a maximal LINPACK benchmark performance of 1.059 PFlop\/s and a theoretical peak performance of 1.3814 PFlop\/s. Annually, 80 percent of Jaguar&amp;#39;s resources are allocated through the U.S Department of Energy's Innovative and Novel Computational Impact on Theory and Experiment (INCITE) program, a competitively selected, peer reviewed process open to researchers from universities, industry, government and non-profit organizations. These allocations address some of the most challenging scientific problems in areas such as climate modeling, renewable energy, materials science, fusion and combustion. In conjunction with NCCS, the Computer Science and Mathematics Division at ORNL performs basic and applied research in HPC, mathematics, and intelligent systems. This talk gives a summary of the HPC research performed at ORNL. It provides details about the Jaguar peta-scale computing resource, an overview of the computational science research carried out using ORNL's computing resources, and a description of various computer science efforts targeting solutions for next-generation HPC systems. This talk also provides information about internship opportunities for MSc students.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Modular Redundancy for Soft-Error Resilience in Large-Scale HPC Systems<\/b>. Invited talk at the <a href=\"http:\/\/www.dagstuhl.de\/en\/program\/calendar\/semhp\/?semnr=09191\" target=\"www.dagstuhl.de\/en\/program\/calendar\/semhp\/?semnr=09191\">Dagstuhl Seminar on Fault Tolerance in High-Performance Computing and Grids<\/a>, Schloss Dagstuhl, Wadern, Germany, May 3-8, 2009. <a href=\"javascript:showAbstract('Recent investigations into resilience of large-scale high-performance computing (HPC) systems showed a continuous trend of decreasing reliability and availability. Newly installed systems have a lower mean-time to failure (MTTF) and a higher mean-time to recover (MTTR) than their predecessors. Modular redundancy is being used in many mission critical systems today to provide for resilience, such as for aerospace and command &amp; control systems. The primary argument against modular redundancy for resilience in HPC has always been that the capability of a HPC system, and respective return on investment, would be significantly reduced. We argue that modular redundancy can significantly increase compute node availability as it removes the impact of scale from single compute node MTTR. We further argue that single compute nodes can be much less reliable, and therefore less expensive, and still be highly available, if their MTTR\/MTTF ratio is maintained.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09modular.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09modular\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Proactive Fault Tolerance Using Preemptive Migration<\/b>. Invited talk at the <a href=\"http:\/\/acet.rdg.ac.uk\/events\/details\/cancun.php\" target=\"acet.rdg.ac.uk\/events\/details\/cancun.php\">3rd Collaborative and Grid Computing Technologies Workshop (CGCTW) 2009<\/a>, Cancun, Mexico, April 22-24, 2009. <a href=\"javascript:showAbstract('The continuing growth in high-performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation HPC systems. The concept of proactive fault tolerance prevents compute node failures from impacting running parallel applications by preemptively migrating application parts away from nodes that are about to fail. This talk presents our past and ongoing efforts in proactive fault resilience for HPC. Presented work includes proactive fault resilience techniques, transparent process- and virtual-machine-level migration, system and application reliability models and analyses, failure prediction, and trade-off models for combining preemptive migration with checkpoint\/restart. All these individual technologies are put into context with a proposed holistic HPC fault resilience framework.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann09proactive2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann09proactive2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resiliency<\/b>. Panel at the <a href=\"http:\/\/www.cs.sandia.gov\/Conferences\/SOS13\" target=\"www.cs.sandia.gov\/Conferences\/SOS13\">13th Workshop on Distributed Supercomputing (SOS) 2009<\/a>, Hilton Head, SC, USA, March 9-12, 2009. <a href=\"?page_id=55#engelmann09resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High-Performance Computing Research at Oak Ridge National Laboratory<\/b>. Invited talk at the Reading Annual Computational Science  Workshop, Reading, United Kingdom, December 8, 2008. <a href=\"javascript:showAbstract('Oak Ridge National Laboratory (ORNL) is the largest energy laboratory in the United States. Its National Center for Computational Sciences (NCCS) provides the most powerful computing resources in the world for open scientific research. Jaguar, a Cray XT5 system at NCCS, is the second HPC system to exceed 1 PFlop\/s (10^15 Floating Point Operations Per Second), and the fastest open science supercomputer in the world. It recently ranked #2 in the Top 500 List of Supercomputer Sites with a maximal LINPACK benchmark performance of 1.059 PFlop\/s and a theoretical peak performance of 1.3814 PFlop\/s. Annually, 80 percent of Jaguar\u2019s resources are allocated through the U.S Department of Energy\u2019s Innovative and Novel Computational Impact on Theory and Experiment (INCITE) program, a competitively selected, peer reviewed process open to researchers from universities, industry, government and non-profit organizations. These allocations address some of the most challenging scientific problems in areas such as climate modeling, renewable energy, materials science, fusion and combustion. In conjunction with NCCS, the Computer Science and Mathematics Division at ORNL performs basic and applied research in HPC, mathematics, and intelligent systems. This talk gives a summary of the HPC research performed at ORNL. It provides details about the Jaguar peta-scale computing resource, an overview of the computational science research carried out using ORNL\u2019s computing resources, and a description of various computer science efforts targeting solutions for next-generation HPC systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Modular Redundancy in HPC Systems: Why, Where, When and How?<\/b>. Invited talk at the 1st HPC Resiliency Summit: Workshop on Resiliency for Petascale HPC 2008, in conjunction with the <a href=\"http:\/\/www.lanl.gov\/conferences\/lacss\/2008\" target=\"www.lanl.gov\/conferences\/lacss\/2008\">1st Los Alamos Computer Science Symposium (LACSS) 2008<\/a>, Santa Fe, NM, USA, October 15, 2008. <a href=\"javascript:showAbstract('The continuing growth in high-performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our past accomplishments, ongoing work, and future plans in the area of high-level RAS for HPC.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08modular.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08modular\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resiliency for High-Performance Computing<\/b>. Invited talk at the <a href=\"http:\/\/acet.rdg.ac.uk\/events\/details\/cancun.php\" target=\"acet.rdg.ac.uk\/events\/details\/cancun.php\">2nd Collaborative and Grid Computing Technologies Workshop (CGCTW) 2008<\/a>, Cancun, Mexico, April 10-12, 2008. <a href=\"javascript:showAbstract('In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation high-performance computing (HPC) systems. One major source of concern are non-recoverable soft errors, i.e., bit flips in memory, cache, registers, and logic. The probability of such errors not only grows with system size, but also with increasing architectural vulnerability caused by employing accelerators and by shrinking nanometer technology. Reactive fault tolerance technologies, such as checkpoint\/restart, are unable to handle high failure rates due to associated overheads, while proactive resiliency technologies, such as preemptive migration, simply fail as random soft errors can&amp;#39;t be predicted. This talk proposes a new, bold direction in resiliency for HPC as it targets resiliency for next-generation extreme-scale HPC systems at the system software level through computational redundancy strategies, i.e., dual- and triple-modular redundancy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08resiliency.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Advanced Fault Tolerance Solutions for High Performance Computing<\/b>. Seminar at the <a href=\"http:\/\/www.laas.fr\" target=\"www.laas.fr\">Laboratoire d&#39;Analyse et d&#8217;Architecture des Syst&eacute;mes<\/a>, <a href=\"http:\/\/www.cnrs.fr\" target=\"www.cnrs.fr\">Centre National de la Recherche Scientifique<\/a>, Toulouse, France, February 11, 2008. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08advanced.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08advanced\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Service-Level High Availability in Parallel and Distributed Systems<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, October 10, 2007. <a href=\"javascript:showAbstract('As service-oriented architectures become more important in parallel and distributed computing systems, individual service instance reliability as well as appropriate service redundancy are essential to increase overall system availability. This talk focuses on redundancy strategies using service-level replication techniques. An overview of existing programming models for service-level high availability is presented and their differences, similarities, advantages, and disadvantages are discussed. Recent advances in providing service-level symmetric active\/active high availability are discussed. While the primary target of the presented research is high availability for service nodes in tightly-coupled extreme-scale high-performance computing (HPC) systems, it is also applicable to loosely-coupled distributed computing scenarios.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07service.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07service\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Advanced Fault Tolerance Solutions for High Performance Computing<\/b>. Invited talk at the <a href=\"http:\/\/www.thaigrid.or.th\/wttc2007\" target=\"www.thaigrid.or.th\/wttc2007\">Workshop on Trends, Technologies and Collaborative Opportunities in High Performance and Grid Computing (WTTC) 2007<\/a>, Khon Kean, Thailand, June 8, 2007. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07advanced2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07advanced2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Advanced Fault Tolerance Solutions for High Performance Computing<\/b>. Invited talk at the <a href=\"http:\/\/www.thaigrid.or.th\/wttc2007\" target=\"www.thaigrid.or.th\/wttc2007\">Workshop on Trends, Technologies and Collaborative Opportunities in High Performance and Grid Computing (WTTC) 2007<\/a>, Bangkok, Thailand, June 4-5, 2007. <a href=\"javascript:showAbstract('The continuing growth in high performance computing (HPC) system scale poses a challenge for system software and scientific applications with respect to reliability, availability and serviceability (RAS). With only very few exceptions, the availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. As a result, sites lower allowable job run times in order to force applications to store intermediate results (checkpoints) as insurance against lost computation time. However, checkpoints themselves waste valuable computation time and resources. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically with the trend towards capability computing, which drives the race for scientific discovery by running applications on the fastest machines available while desiring significant amounts of time (weeks and months) without interruption. These machines must be able to run in the event of frequent interrupts in such a manner that the capability is not severely degraded. Thus, research and development of scalable RAS technologies is paramount to the success of future extreme-scale systems. This talk summarizes our accomplishments in the area of high-level RAS for HPC, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07advanced.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07advanced\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Operating System Research at ORNL: System-level Virtualization<\/b>. Seminar at the <a href=\"http:\/\/www.gup.uni-linz.ac.at\" target=\"www.gup.uni-linz.ac.at\">Institute of Graphics and Parallel Processing<\/a>, <a href=\"http:\/\/www.uni-linz.ac.at\" target=\"www.uni-linz.ac.at\">Johannes Kepler University<\/a>, Linz, Austria, April 10, 2007. <a href=\"javascript:showAbstract('The emergence of virtualization enabled hardware, such as the latest generation AMD and Intel processors, has raised significant interest in High Performance Computing (HPC) community. In particular, system-level virtualization provides an opportunity to advance the design and development of operating systems, programming environments, administration practices, and resource management tools. This leads to some potential research topics for HPC, such as failure tolerance, system management, and solutions for application porting to new HPC platforms. This talk will present an overview of the research in System-level Virtualization taking place by the Systems Research Team in the Computer Science Research Group at Oak Ridge National Laboratory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07operating.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07operating\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Towards High Availability for High-Performance Computing System Services: Accomplishments and Limitations<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, March 14, 2007. <a href=\"javascript:showAbstract('During the last several years, our teams at Oak Ridge National Laboratory, Louisiana Tech University, and Tennessee Technological University focused on efficient redundancy strategies for head and service nodes of high-performance computing (HPC) systems in order to pave the way for high availability (HA) in HPC. These nodes typically run critical HPC system services, like job and resource management, and represent single points of failure and control for an entire HPC system. The overarching goal of our research is to provide high-level reliability, availability, and serviceability (RAS) for HPC systems by combining HA and HPC technology. This talk summarizes our accomplishments, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07towards.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, June 9, 2006. <a href=\"javascript:showAbstract('A major concern in exploiting ultra-scale architectures for scientific high-end computing (HEC) with tens to hundreds of thousands of processors, such as the IBM Blue Gene\/L and the Cray X1, is the potential inability to identify problems and take preemptive action before a failure impacts a running job. In fact, in systems of this scale, predictions estimate the mean time to interrupt in terms of hours. Current solutions for fault-tolerance in HEC focus on dealing with the result of a failure. However, most are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services (e.g. MPI) or even of the entire machine. High availability (HA) computing strives to avoid the problems of unexpected failures through preemptive measures. There are various techniques to implement high availability. In contrast to active\/hot-standby high availability with its fail-over model, active\/active high availability with its virtual synchrony model is superior in many areas including scalability, throughput, availability and responsiveness. However, it is significantly more complex. The overall goal of our research is to expand today`s effort in HA for HEC, so that systems that have the ability to hot-swap hardware components can be kept alive by an OS runtime environment that understands the concept of dynamic system configuration. This talk will present an overview of recent research at Oak Ridge National Laboratory in high availability solutions for ultra-scale scientific high-end computing.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott and Christian Engelmann. <b>Advancing Reliability, Availability and Serviceability for High-Performance Computing<\/b>. Seminar at the <a href=\"http:\/\/www.gup.uni-linz.ac.at\" target=\"www.gup.uni-linz.ac.at\">Institute of Graphics and Parallel Processing<\/a>, <a href=\"http:\/\/www.uni-linz.ac.at\" target=\"www.uni-linz.ac.at\">Johannes Kepler University<\/a>, Linz, Austria, April 19, 2006. <a href=\"javascript:showAbstract('Today\u2019s high performance computing systems have several reliability deficiencies resulting in noticeable availability and serviceability issues. For example, head and service nodes represent a single point of failure and control for an entire system as they render it inaccessible and unmanageable in case of a failure until repair, causing a significant downtime. Furthermore, current solutions for fault-tolerance focus on dealing with the result of a failure. However, most are unable to transparently mask runtime system configuration changes caused by failures and require a complete restart of essential system services, such as MPI, in case of a failure. High availability computing strives to avoid the problems of unexpected failures through preemptive measures. The overall goal of our research is to expand today\u2019s effort in high availability for high-performance computing, so that systems can be kept alive by an OS runtime environment that understands the concepts of dynamic system configuration and degraded operation mode. This talk will present an overview of recent research performed at Oak Ridge National Laboratory in collaboration with Louisiana Tech University, North Carolina State University and the University of Reading in developing core technologies and proof-of-concept prototypes that improve the overall reliability, availability and serviceability of high-performance computing systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott06advancing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#scott06advancing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, October 18, 2005. <a href=\"javascript:showAbstract('A major concern in exploiting ultra-scale architectures for scientific high-end computing (HEC) with tens to hundreds of thousands of processors, such as the IBM Blue Gene\/L and the Cray X1, is the potential inability to identify problems and take preemptive action before a failure impacts a running job. In fact, in systems of this scale, predictions estimate the mean time to interrupt in terms of hours. Current solutions for fault-tolerance in HEC focus on dealing with the result of a failure. However, most are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services (e.g. MPI) or even of the entire machine. High availability (HA) computing strives to avoid the problems of unexpected failures through preemptive measures. There are various techniques to implement high availability. In contrast to active\/hot-standby high availability with its fail-over model, active\/active high availability with its virtual synchrony model is superior in many areas including scalability, throughput, availability and responsiveness. However, it is significantly more complex. The overall goal of our research is to expand today`s effort in HA for HEC, so that systems that have the ability to hot-swap hardware components can be kept alive by an OS runtime environment that understands the concept of dynamic system configuration. This talk will present an overview of recent research at Oak Ridge National Laboratory in high availability solutions for ultra-scale scientific high-end computing.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05high4.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05high4\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. Seminar at the <a href=\"http:\/\/www.uncfsu.edu\/macsc\" target=\"www.uncfsu.edu\/macsc\">Department of Mathematics and Computer Science<\/a>, <a href=\"http:\/\/www.uncfsu.edu\" target=\"www.uncfsu.edu\">Fayetteville State University<\/a>, Fayetteville, NC, USA, September 26, 2005. <a href=\"javascript:showAbstract('A major concern in exploiting ultra-scale architectures for scientific high-end computing (HEC) with tens to hundreds of thousands of processors, such as the IBM Blue Gene\/L and the Cray X1, is the potential inability to identify problems and take preemptive action before a failure impacts a running job. In fact, in systems of this scale, predictions estimate the mean time to interrupt in terms of hours. Current solutions for fault-tolerance in HEC focus on dealing with the result of a failure. However, most are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services (e.g. MPI) or even of the entire machine. High availability (HA) computing strives to avoid the problems of unexpected failures through preemptive measures. There are various techniques to implement high availability. In contrast to active\/hot-standby high availability with its fail-over model, active\/active high availability with its virtual synchrony model is superior in many areas including scalability, throughput, availability and responsiveness. However, it is significantly more complex. The overall goal of our research is to expand today\u2019s effort in HA for HEC, so that systems that have the ability to hot-swap hardware components can be kept alive by an OS runtime environment that understands the concept of dynamic system configuration. This talk will present an overview of recent research at Oak Ridge National Laboratory in fault tolerance and high availability solutions for ultra-scale scientific high-end computing.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05high3.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05high3\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. Seminar at the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, Reading, United Kingdom, May 13, 2005. <a href=\"javascript:showAbstract('A major concern in exploiting ultra-scale architectures for scientific high-end computing (HEC) with tens to hundreds of thousands of processors, such as the IBM Blue Gene\/L and the Cray X1, is the potential inability to identify problems and take preemptive action before a failure impacts a running job. In fact, in systems of this scale, predictions estimate the mean time to interrupt in terms of hours. Current solutions for fault-tolerance in HEC focus on dealing with the result of a failure. However, most are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services (e.g. MPI) or even of the entire machine. High availability (HA) computing strives to avoid the problems of unexpected failures through preemptive measures. There are various techniques to implement high availability. In contrast to active\/hot-standby high availability with its fail-over model, active\/active high availability with its virtual synchrony model is superior in many areas including scalability, throughput, availability and responsiveness. However, it is significantly more complex. The overall goal of our research is to expand today\u2019s effort in HA for HEC, so that systems that have the ability to hot-swap hardware components can be kept alive by an OS runtime environment that understands the concept of dynamic system configuration. This talk will present an overview of recent research at Oak Ridge National Laboratory in fault-tolerant heterogeneous metacomputing, advanced super-scalable algorithms and high availability system software for ultra-scale scientific high-end computing.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05high2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05high2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. Seminar at the <a href=\"http:\/\/cenit.latech.edu\" target=\"cenit.latech.edu\">Center for Entrepreneurship and Information Technology<\/a>, <a href=\"http:\/\/www.latech.edu\" target=\"www.latech.edu\">Louisiana Tech University<\/a>, Ruston, LA, USA, April 15, 2005. <a href=\"javascript:showAbstract('A major concern in exploiting ultra-scale architectures for scientific high-end computing (HEC) with tens to hundreds of thousands of processors is the potential inability to identify problems and take preemptive action before a failure impacts a running job. In fact, in systems of this scale, predictions estimate the mean time to interrupt in terms of hours. Current solutions for fault-tolerance in HEC focus on dealing with the result of a failure. However, most are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services (e.g. MPI) or even of the entire machine. High availability (HA) computing strives to avoid the problems of unexpected failures through preemptive measures. There are various techniques to implement high availability. In contrast to active\/hot-standby high availability with its fail-over model, active\/active high availability with its virtual synchrony model is superior in many areas including scalability, throughput, availability and responsiveness. However, it is significantly more complex. The overall goal of this research is to expand today\u2019s effort in HA for HEC, so that systems that have the ability to hot-swap hardware components can be kept alive by an OS runtime environment that understands the concept of dynamic system configuration. With the aim of addressing the future challenges of high availability in ultra-scale HEC, this project intends to develop a proof-of-concept implementation of an active\/active high availability system software framework.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05high1.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05high1\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Diskless Checkpointing on Super-scale Architectures &#8211; Applied to the Fast Fourier Transform<\/b>. Invited talk at the <a href=\"http:\/\/www.siam.org\/meetings\/pp04\" target=\"www.siam.org\/meetings\/pp04\">11th SIAM Conference on Parallel Processing for Scientific Computing (SIAM PP) 2004<\/a>, San Francisco, CA, USA, February 25, 2004. <a href=\"javascript:showAbstract('This talk discusses the issue of fault-tolerance in distributed computer systems with tens or hundreds of thousands of diskless processor units. Such systems, like the IBM Blue Gene\/L, are predicted to be deployed in the next five to ten years. Since a 100,000-processor system is going to be less reliable, scientific applications need to be able to recover from occurring failures more efficiently. In this paper, we adapt the present technique of diskless checkpointing to such huge distributed systems in order to equip existing scientific algorithms with super-scalable fault-tolerance. First, we discuss the method of diskless checkpointing, then we adapt this technique to super-scale architectures and finally we present results from an implementation of the Fast Fourier Transform that uses the adapted technique to achieve super-scale fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann04diskless.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann04diskless\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Super-scalable Algorithms &#8211; Next Generation Supercomputing on 100,000 and more Processors<\/b>. Seminar at the <a href=\"http:\/\/www.csm.ornl.gov\" target=\"www.csm.ornl.gov\">Computer Science and Mathematics Division<\/a>, <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\">Oak Ridge National Laboratory<\/a>, Oak Ridge, TN, USA, January 29, 2004. <a href=\"javascript:showAbstract('This talk discusses recent research into the issues and potential problems of algorithm scalability and fault-tolerance on next-generation high-performance computer systems with tens and even hundreds of thousands of processors. Such massively parallel computers, like the IBM Blue Gene\/L, are going to be deployed in the next five to ten years and existing deficiencies in scalability and fault-tolerance need to be addressed soon. Scientific algorithms have shown poor scalability on 10,000-processor systems that exist today. Furthermore, future systems will be less reliable due to the large number of components. Super-scalable algorithms, which have the properties of scale invariance and natural fault-tolerance, are able to get the correct answer despite multiple task failures and without checkpointing. We will show that such algorithms exist for a wide variety of problems, such as finite difference, finite element, multigrid and global maximum. Despite these findings, traditional algorithms may still be preferred due to their known behavior, or simply because a super-scalable algorithm does not exist or is hard to find for a particular problem. In this case, we propose a peer-to-peer diskless checkpointing algorithm that can provide scale invariant fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann04superscalable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann04superscalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Distributed Peer-to-Peer Control for Harness<\/b>. Seminar at the <a href=\"http:\/\/www.csc.ncsu.edu\" target=\"www.csc.ncsu.edu\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.ncsu.edu\" target=\"www.ncsu.edu\">North Carolina State University<\/a>, Raleigh, NC, USA, February 11, 2004. <a href=\"javascript:showAbstract('Harness is an adaptable fault-tolerant virtual machine environment for next-generation heterogeneous distributed computing developed as a follow on to PVM. It additionally enables the assembly of applications from plug-ins and provides fault-tolerance. This work describes the distributed control, which manages global state replication to ensure a high-availability of service. Group communication services achieve an agreement on an initial global state and a linear history of global state changes at all members of the distributed virtual machine. This global state is replicated to all members to easily recover from single, multiple and cascaded faults. A peer-to-peer ring network architecture and tunable multi-point failure conditions provide heterogeneity and scalability. Finally, the integration of the distributed control into the multi-threaded kernel architecture of Harness offers a fault-tolerant global state database service for plug-ins and applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann03distributed.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann03distributed\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"students\"><\/a>Co-advised Theses<\/h4>\n<ol>\n<li>Ian S. Jones. <b>Simulation of Large Scale Architectures on High Performance Computers<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, October 22, 2010. Thesis research performed at Oak Ridge National Laboratory. Advisors: Prof. Vassil N. Alexandrov (University of Reading); Christian Engelmann (Oak Ridge National Laboratory); George Bosilca (University of Tennessee, Knoxville). <a href=\"javascript:showAbstract('Powerful supercomputers often need to be simulated for the purposes of testing the scalability of various applications. This thesis endeavours to further develop the existing simulator, XSIM, and implement the functionality to simulate real-world networks and the latency which might be encountered by messages travelling through that network. The upgraded simulator will then be tested at the Oak Ridge National Laboratory. The work completed herein should provide a solid foundation for further improvements to XSIM; it simulates a variety of basic network topologies, calculating the shortest path for any given message and generates a transmission time.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/jones10simulation.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/jones10simulation.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#jones10simulation\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen B&ouml;hm. <b>Development of a RAS Framework for HPC Environments: Realtime Data Reduction of Monitoring Data<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, March 12, 2010. Thesis research performed at Oak Ridge National Laboratory. Advisors: Prof. Vassil N. Alexandrov (University of Reading); Christian Engelmann (Oak Ridge National Laboratory); George Bosilca (University of Tennessee, Knoxville). <a href=\"javascript:showAbstract('The advancements of high-performance computing (HPC) systems in the last decades lead to more and more complex systems containing thousands or tens-of-thousands computing systems that are working together. While the computational performance of these systems increased dramaticaly in the last years the I\/O subsystems have not gained such a significant improvement. With increasing nummbers of hardware components in the next generation HPC systems maintaining the relaiability of such systems becomes more and more difficult since the probability of hardware failures is increasing with the number of components. The capacities of traditional reactive fault tolerance technologies are exceeded by the development of next generation systems and alternatives have to be found. This paper discusses a monitoring system that is using data reduction techniques to decrease the amount of the collected data. The system is part of a proactive fault tolerance system that may challenge the reliability problems of exascale HPC systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm10development.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/boehm10development.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#boehm10development\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Frank Lauer. <b>Simulation of Advanced Large-Scale HPC Architectures<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, March 12, 2010. Thesis research performed at Oak Ridge National Laboratory. Advisors: Prof. Vassil N. Alexandrov (University of Reading); Christian Engelmann (Oak Ridge National Laboratory); George Bosilca (University of Tennessee, Knoxville). <a href=\"javascript:showAbstract('The rapid development of massive parallel systems in the high- performance computing (HPC) area requires efficient scalability of applications. The next generation&amp;#39;s design of supercomputers is today not certain in terms of what will be the computational, memory and I\/O capabilities. However it is most certain that they become even more parallel. Getting the most performance from these machines in not only a matter of hardware, it is also an issue of programming design. Therefore, it has to be a co-development. However, how to test algorithm's on machines which are not existing today. To address the programming issues in terms of scalability and fault tolerance for the next generation, this projects aim is to design and develop a simulator based on parallel discrete event simulation (PDES) for applications using MPI communication. Some of the fastest supercomputers in the world already interconnecting &amp;#36;10^5 cores together to catch up the simulator will be able to simulate at least 10^7 virtual processes.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/lauer10simulation.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/lauer10simulation.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#lauer10simulation\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Antonina Litvinova. <b>RAS Framework Engine Prototype<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, September 22, 2009. Thesis research performed at Oak Ridge National Laboratory. Advisors: Prof. Vassil N. Alexandrov (University of Reading); Christian Engelmann (Oak Ridge National Laboratory); George Bosilca (University of Tennessee, Knoxville). <a href=\"javascript:showAbstract('Extreme high performance computing (HPC) systems constantly increase in scale from a few thousands of processors cores to thousands of thousands of processors cores and beyond. However their system mean-time to interrupt decreases according. The current approach of fault tolerance in HPC is checkpoint\/restart, i.e. a method based on recovery from experienced failures. However checkpoint\/restart cannot deal with errors in the same efficient way anymore, because of HPC systems modification. For example, increasing error rates, increasing aggregate memory, and not proportionally increasing input\/output capabilities. The recently introduced concept is proactive fault tolerance which avoids experiencing failures through preventative measures. Proactive fault tolerance uses migration which is an emerging technology that prevents failures on HPC systems by migrating applications or application parts away from a node that is deteriorating to a spare node. This thesis discusses work conducted at ORNL to develop a Proactive Fault Tolerance Framework Engine Prototype for HPC systems with high reliability, availability and serviceability. The prototype performs environmental system monitoring, system event logging, parallel job monitoring and system resource monitoring in order to analyse HPC system reliability and to perform fault avoidance through a migration.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/litvinova09ras.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/litvinova09ras.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#litvinova09ras\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Bj&ouml;rn K&ouml;nning. <b>Virtualized Environments for the Harness Workbench<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, March 14, 2007. Thesis research performed at Oak Ridge National Laboratory. Advisors: Prof. Vassil N. Alexandrov (University of Reading); Christian Engelmann (Oak Ridge National Laboratory). <a href=\"javascript:showAbstract('The expanded use of computational sciences today leads to a significant need of high performance computing systems. High performance computing is currently undergoing vigorous revival, and multiple efforts are underway to develop much faster computing systems in the near future. New software tools are required for the efficient use of petascale computing systems. With the new Harness Workbench Project the Oak Ridge National Laboratory intends to develop an appropriate development and runtime environment for high performance computing platforms. This dissertation project is part of the Harness Workbench Project, and deals with the development of a concept for virtualised environments and various approaches to create and describe them. The developed virtualisation approach is based on the \\verb|chroot| mechanism and uses platform-independent environment descriptions. File structures and environment variables are emulated to provide the portability of computational software over diverse high performance computing platforms. Security measures and sandbox characteristic are integrable.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/koenning07virtualized.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/koenning07virtualized.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#koenning07virtualized\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Matthias Weber. <b>High Availability for the Lustre File System<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, March 14, 2007. Thesis research performed at Oak Ridge National Laboratory. Double diploma in conjunction with the <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Department of Engineering I<\/a>, <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Technical College for Engineering and Economics (FHTW) Berlin<\/a>, Germany. Advisors: Prof. Vassil N. Alexandrov (University of Reading); Christian Engelmann (Oak Ridge National Laboratory). <a href=\"javascript:showAbstract('With the growing importance of high performance computing and, more importantly, the fast growing size of sophisticated high performance computing systems, research in the area of high availability is essential to meet the needs to sustain the current growth. This Master thesis project aims to improve the availability of Lustre. Major concern of this project is the metadata server of the file system. The metadata server of Lustre suffers from the last single point of failure in the file system. To overcome this single point of failure an active\/active high availability approach is introduced. The new file system design with multiple MDS nodes running in virtual synchrony leads to a significant increase of availability. Two prototype implementations aim to show how the proposed system design and its new realized form of symmetric active\/active high availability can be accomplished in practice. The results of this work point out the difficulties in adapting the file system to the active\/active high availability design. Tests identify not achieved functionality and show performance problems of the proposed solution. The findings of this dissertation may be used for further work on high availability for distributed file systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/weber07high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/weber07high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#weber07high\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Ronald Baumann. <b>Design and Development of Prototype Components for the Harness High-Performance Computing Workbench<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, March 6, 2006. Thesis research performed at Oak Ridge National Laboratory. Double diploma in conjunction with the <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Department of Engineering I<\/a>, <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Technical College for Engineering and Economics (FHTW) Berlin<\/a>, Germany. Advisors: Prof. Vassil N. Alexandrov (University of Reading); George A. (Al) Geist and Christian  Engelmann (Oak Ridge National Laboratory). <a href=\"javascript:showAbstract('This master thesis examines plug-in technology, especially the new field of parallel plug-ins. Plug-ins are popular because they extend the capabilities of software packages such as browsers and Photoshop, and allow an individual user to add new functionality. Parallel plug-ins also provide the above capabilities to a distributed set of resources, i.e., a plug-in now becomes a set of coordinating plug-ins. Second, the set of plugins may be heterogeneous either in function or because the underlying resources are heterogeneous. This new dimension of complexity provides a rich research space which is explored in this thesis. Experiences are collected and presented as parallel plug-in paradigms and concepts. The Harness framework was used in this project, in particular the plugin manager and available communication capabilities. Plug-ins provide methods for users to extend Harness according to their requirements. The result of this thesis is a parallel plug-in paradigm and template for Harness. Users of the Harness environment will be able to design and implement their applications in the form of parallel plug-ins easier and faster by using the paradigm resulting from this project. Prototypes were implemented which handle different aspects of parallel plug-ins. Parallel plug-in configurations were tested on an appropriate number of Harness kernels, including available communication and error-handling capabilities. Furthermore, research was done in the area of fault tolerance while parallel plug-ins are (un)loaded, as well as while a task is performed.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/baumann06design.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/baumann06design.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#baumann06design\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Kai Uhlemann. <b>High Availability for High-End Scientific Computing<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, March 6, 2006. Thesis research performed at Oak Ridge National Laboratory. Double diploma in conjunction with the <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Department of Engineering I<\/a>, <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Technical College for Engineering and Economics (FHTW) Berlin<\/a>, Germany. Advisors: Prof. Vassil N. Alexandrov (University of Reading); George A. (Al) Geist and  Christian Engelmann (Oak Ridge National Laboratory). <a href=\"javascript:showAbstract('With the growing interest and popularity in high performance cluster computing and, more importantly, the fast growing size of compute clusters, research in the area of high availability is essential to meet the needs to sustain the current growth. This Master thesis project introduces a new approach for high availability focusing on the head node of a cluster system. This projects focus is on providing high availability to the job scheduler service, which is the most vital part of the traditional Beowulf-style cluster architecture. This research seeks to add high availability to the job scheduler service and resource management system, typically running on the head node, leading to a significant increase of availability for cluster computing. Also, this software project takes advantage of the virtual synchrony paradigm to achieve active\/active replication, the highest form of high availability. A proof-of-concept implementation shows how high availability can be designed in software and what results can be expected of such a system. The results may be reused for future or existing projects to further improve and extent the high availability of compute clusters.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/uhlemann06high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/uhlemann06high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#uhlemann06high\" ><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>\n<a name=\"theses\"><\/a>Theses<\/h4>\n<ol>\n<li>Christian Engelmann. <b>Symmetric Active\/Active High Availability for High-Performance Computing System Services<\/b>. PhD thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, December 8, 2008. Thesis research performed at Oak Ridge National Laboratory. Advisor: Prof. Vassil N. Alexandrov (University of Reading). <a href=\"javascript:showAbstract('In order to address anticipated high failure rates, reliability, availability and serviceability have become an urgent priority for next-generation high-performance computing (HPC) systems. This thesis aims to pave the way for highly available HPC systems by focusing on their most critical components and by reinforcing them with appropriate high availability solutions. Service components, such as head and service nodes, are the Achilles heel of a HPC system. A failure typically results in a complete system-wide outage. This thesis targets efficient software state replication mechanisms for service component redundancy to achieve high availability as well as high performance. Its methodology relies on defining a modern theoretical foundation for providing service-level high availability, identifying availability deficiencies of HPC systems, and comparing various service-level high availability methods. This thesis showcases several developed proof-of-concept prototypes providing high availability for services running on HPC head and service nodes using the symmetric active\/active replication method, i.e., state-machine replication, to complement prior work in this area using active\/standby and asymmetric active\/active configurations. Presented contributions include a generic taxonomy for service high availability, an insight into availability deficiencies of HPC systems, and a unified definition of service-level high availability methods. Further contributions encompass a fully functional symmetric active\/active high availability prototype for a HPC job and resource management service that does not require modification of service, a fully functional symmetric active\/active high availability prototype for a HPC parallel file system metadata service that offers high performance, and two preliminary prototypes for a transparent symmetric active\/active replication software framework for client-service and dependent service scenarios that hide the replication infrastructure from clients and services. Assuming a mean-time to failure of 5,000 hours for a head or service node, all presented prototypes improve service availability from 99.285% to 99.995% in a two-node system, and to 99.99996% with three nodes.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08symmetric3.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann08symmetric3.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08symmetric3\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Distributed Peer-to-Peer Control for Harness<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK, July 7, 2001. Thesis research performed at Oak Ridge National Laboratory. Double diploma in conjunction with the <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Department of Engineering I<\/a>, <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Technical College for Engineering and Economics (FHTW) Berlin<\/a>, Germany. Advisors: Prof. Vassil N. Alexandrov (University of Reading); George A. (Al) Geist (Oak Ridge National Laboratory). <a href=\"javascript:showAbstract('Parallel processing, the method of cutting down a large computational problem into many small tasks which are solved in parallel, is a field of increasing importance in science. Cost-effective, flexible and efficient simulations of mathematical models of physical, chemical or biological real-world problems are replacing the traditional experimental research. Current software solutions for parallel and scientific computation, like Parallel Virtual Machine and Message Passing Interface, have limitations in handling faults and failures, in utilizing heterogeneous and dynamically changing communication structures, and in enabling migrating or cooperative applications. The current research in heterogeneous adaptable reconfigurable networked systems (Harness) aims to produce the next generation of software solutions for distributed computing. A high-available and light-weighted distributed virtual machine service provides an encapsulation of a few hundred to a few thousand physical machines in a virtual heterogeneous large scale cluster. A high availability of a service in distributed systems can be achieved by replication of the service state on multiple server processes. If one ore more server processes fails, the surviving ones continue to provide the service because they know the state. Since every member of a distributed virtual machine is part of the distributed virtual machine service state and is able to change this state, a distributed control is needed to replicate the state and maintain its consistency. This distributed control manages state changes as well as the state-replication and the detection of and recovery from faults and failures of server processes. This work analyzes system architectures currently used in heterogeneous distributed computing by defining terms, conditions and assumptions. It shows that such systems are asynchronous and may use partially synchronous communication to detect and to distinguish different classes of faults and failures. It describes how a high availability of a large scale distributed service on a huge number of servers residing on different geographical locations can be realized. Asynchronous group communication services, such as Reliable Broadcast, Atomic Broadcast, Distributed Agreement and Membership, are analyzed to develop linear scalable algorithms in an unidirectional and in a bidirectional connected asynchronous peer-to-peer ring architecture. A Transaction Control group communication service is introduced as state-replication service. The system analysis distinguishes different types of distributed systems, where active transactions execute state changes using non-replicated data of one or more servers and inactive transactions report state changes using replicated data only. It is applicable for passive fault-tolerant distributed databases as well as for active fault-tolerant distributed control mechanisms. No control token is used and time stamps are avoided, so that all members of a server group have equal responsibilities and are independent from the system time. A prototype which implements the most complicated Transaction Control algorithm is realized due to the complexity of the distributed system and the early development stage of the introduced algorithms. The prototype is used to obtain practical experience with the state-replication algorithm.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann01distributed.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann01distributed.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann01distributed\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Distributed Peer-to-Peer Control for Harness<\/b>. Master&#8217;s thesis, <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Department of Engineering I<\/a>, <a href=\"http:\/\/www.f1.fhtw-berlin.de\" target=\"www.f1.fhtw-berlin.de\">Technical College for Engineering and Economics (FHTW) Berlin<\/a>, Germany, February 23, 2001. Thesis research performed at Oak Ridge National Laboratory. Double diploma in conjunction with the <a href=\"http:\/\/www.cs.reading.ac.uk\" target=\"www.cs.reading.ac.uk\">Department of Computer Science<\/a>, <a href=\"http:\/\/www.reading.ac.uk\" target=\"www.reading.ac.uk\">University of Reading<\/a>, UK. Advisors: Prof. Uwe Metzler (Technical College for Engineering and Economics (FHTW) Berlin); George A. (Al) Geist (Oak Ridge National Laboratory). <a href=\"javascript:showAbstract('Parallel processing, the method of cutting down a large computational problem into many small tasks which are solved in parallel, is a field of increasing importance in science. Cost-effective, flexible and efficient simulations of mathematical models of physical, chemical or biological real-world problems are replacing the traditional experimental research. Current software solutions for parallel and scientific computation, like Parallel Virtual Machine and Message Passing Interface, have limitations in handling faults and failures, in utilizing heterogeneous and dynamically changing communication structures, and in enabling migrating or cooperative applications. The current research in heterogeneous adaptable reconfigurable networked systems (Harness) aims to produce the next generation of software solutions for distributed computing. A high-available and light-weighted distributed virtual machine service provides an encapsulation of a few hundred to a few thousand physical machines in a virtual heterogeneous large scale cluster. A high availability of a service in distributed systems can be achieved by replication of the service state on multiple server processes. If one ore more server processes fails, the surviving ones continue to provide the service because they know the state. Since every member of a distributed virtual machine is part of the distributed virtual machine service state and is able to change this state, a distributed control is needed to replicate the state and maintain its consistency. This distributed control manages state changes as well as the state-replication and the detection of and recovery from faults and failures of server processes. This work analyzes system architectures currently used in heterogeneous distributed computing by defining terms, conditions and assumptions. It shows that such systems are asynchronous and may use partially synchronous communication to detect and to distinguish different classes of faults and failures. It describes how a high availability of a large scale distributed service on a huge number of servers residing on different geographical locations can be realized. Asynchronous group communication services, such as Reliable Broadcast, Atomic Broadcast, Distributed Agreement and Membership, are analyzed to develop linear scalable algorithms in an unidirectional and in a bidirectional connected asynchronous peer-to-peer ring architecture. A Transaction Control group communication service is introduced as state-replication service. The system analysis distinguishes different types of distributed systems, where active transactions execute state changes using non-replicated data of one or more servers and inactive transactions report state changes using replicated data only. It is applicable for passive fault-tolerant distributed databases as well as for active fault-tolerant distributed control mechanisms. No control token is used and time stamps are avoided, so that all members of a server group have equal responsibilities and are independent from the system time. A prototype which implements the most complicated Transaction Control algorithm is realized due to the complexity of the distributed system and the early development stage of the introduced algorithms. The prototype is used to obtain practical experience with the state-replication algorithm.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann01distributed2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann01distributed2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann01distributed2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Peer-reviewed Journal Papers Emmanuel Agullo, Mirco Altenbernd, Hartwig Anzt, Leonardo Bautista-Gomez, Tommaso Benacchio, Luca Bonaventura, Hans-Joachim Bungartz, Sanjay Chatterjee, Florina M. Ciorba, Nathan DeBardeleben, Daniel Drzisga, Sebastian Eibl, Christian Engelmann, Wilfried N. Gansterer, Luc Giraud, Dominik G&ouml;ddeke, Marco Heisig, Fabienne J&eacute;z&eacute;quel, Nils Kohl, Xiaoye Sherry Li, Romain Lion, Miriam Mehl, Paul Mycek, Michael Obersteiner, Enrique&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-16","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=16"}],"version-history":[{"count":55,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16\/revisions"}],"predecessor-version":[{"id":1448,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16\/revisions\/1448"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=16"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}