{"id":18,"date":"2025-05-12T08:00:45","date_gmt":"2025-05-12T08:00:45","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=18"},"modified":"2025-05-13T01:51:15","modified_gmt":"2025-05-13T01:51:15","slug":"peer-reviewed-journal-papers","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=18","title":{"rendered":"Peer-Reviewed Journal Papers"},"content":{"rendered":"<ol>\n<li>Emmanuel Agullo, Mirco Altenbernd, Hartwig Anzt, Leonardo Bautista-Gomez, Tommaso Benacchio, Luca Bonaventura, Hans-Joachim Bungartz, Sanjay Chatterjee, Florina M. Ciorba, Nathan DeBardeleben, Daniel Drzisga, Sebastian Eibl, Christian Engelmann, Wilfried N. Gansterer, Luc Giraud, Dominik G&ouml;ddeke, Marco Heisig, Fabienne J&eacute;z&eacute;quel, Nils Kohl, Xiaoye Sherry Li, Romain Lion, Miriam Mehl, Paul Mycek, Michael Obersteiner, Enrique S. Quintana-Ort&iacute;, Francesco Rizzi, Ulrich R&uuml;de, Martin Schulz, Fred Fung, Robert Speck, Linda Stals, Keita Teranishi, Samuel Thibault, Dominik Th&ouml;nnes, Andreas Wagner, and Barbara Wohlmuth. <b>Resiliency in Numerical Algorithm Design for Extreme Scale Simulations<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 36, number 2, pages 251-285, March 1, 2022. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/10943420211055188\" target=\"publication\">10.1177\/10943420211055188<\/a>. <a href=\"javascript:showAbstract('This work is based on the seminar titled &amp;#39;Resiliency in Numerical Algorithm Design for Extreme Scale Simulations' held March 1-6, 2020 at Schloss Dagstuhl, that was attended by all the writers. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 hours on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 1023 floating- point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications, and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/agullo22resiliency.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#agullo22resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar, Saurabh Gupta, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, and Devesh Tiwari. <b>Study of Interconnect Errors, Network Congestion, and Applications Characteristics for Throttle Prediction on a Large Scale HPC System<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 153, pages 29-43, July 1, 2021. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2021.03.001\" target=\"publication\">10.1016\/j.jpdc.2021.03.001<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of peta ops. The performance of these systems depends on various components like processors, memory, and interconnect. Among  all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar21study.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar21study\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Amogh Katti, Giuseppe Di Fatta, Thomas Naughton, and Christian Engelmann. <b>Epidemic Failure Detection and Consensus for Extreme Parallelism<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 32, number 5, pages 729-743, September 1, 2018. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/1094342017690910\" target=\"publication\">10.1177\/1094342017690910<\/a>. <a href=\"javascript:showAbstract('Future extreme-scale high-performance computing systems will be required to work under frequent component failures. The MPI Forum&amp;#39;s User Level Failure Mitigation proposal has introduced an operation, MPI Comm shrink, to synchronize the alive processes on the list of failed processes, so that applications can continue to execute even in the presence of failures by adopting algorithm-based fault tolerance techniques. This MPI Comm shrink operation requires a failure detection and consensus algorithm. This paper presents three novel failure detection and consensus algorithms using Gossiping. The proposed algorithms were implemented and tested using the Extreme-scale Simulator. The results show that in all algorithms the number of Gossip cycles to achieve global consensus scales logarithmically with system size. The second algorithm also shows better scalability in terms of memory and network bandwidth usage and a perfect synchronization in achieving global consensus. The third approach is a three-phase distributed failure detection and consensus algorithm and provides consistency guarantees even in very large and extreme-scale systems while at the same time being memory and bandwidth efficient.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/katti18epidemic.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#katti18epidemic\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale<\/b>. <i><a href=\"http:\/\/superfri.org\/superfri\" target=\"superfri.org\/superfri\">Journal of Supercomputing Frontiers and Innovations (JSFI)<\/a><\/i>, volume 4, number 3, pages 4-42, October 1, 2017. <a href=\"http:\/\/www.susu.ru\/en\" target=\"www.susu.ru\/en\">South Ural State University Chelyabinsk, Russia<\/a>. ISSN 2409-6008. DOI <a href=\"http:\/\/dx.doi.org\/10.14529\/jsfi170301\" target=\"publication\">10.14529\/jsfi170301<\/a>. <a href=\"javascript:showAbstract('Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires management of various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in HPC systems future systems are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. As a result, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods and metrics to investigate and evaluate resilience holistically in HPC systems that consider impact scope, handling coverage, and performance &amp; power efficiency across the system stack. Additionally, few of the current approaches are portable to newer architectures and software environments that will be deployed on future systems. In this paper, we develop a structured approach to the management of HPC resilience using the concept of resilience-based design patterns. A design pattern is a general repeatable solution to a commonly occurring problem. We identify the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. Each established solution is described in the form of a pattern that addresses concrete problems in the design of resilient systems. The complete catalog of resilience design patterns provides designers with reusable design elements. We also define a framework that enhances a designer&amp;#39;s understanding of the important constraints and opportunities for the design patterns to be implemented and deployed at various layers of the system stack. This design framework may be used to establish mechanisms and interfaces to coordinate flexible fault management across hardware and software components. The framework also supports optimization of the cost-benefit trade-offs among performance, resilience, and power consumption. The overall goal of this work is to enable a systematic methodology for the design and evaluation of resilience technologies in extreme-scale HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17resilience.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hukerikar17resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A New Deadlock Resolution Protocol and Message Matching Algorithm for the Extreme-scale Simulator<\/b>. <i><a href=\"http:\/\/onlinelibrary.wiley.com\/journal\/10.1002\/(ISSN)1532-0634\" target=\"onlinelibrary.wiley.com\/journal\/10.1002\/(ISSN)1532-0634\">Concurrency and Computation: Practice and Experience<\/a><\/i>, volume 28, number 12, pages 3369-3389, August 1, 2016. <a href=\"http:\/\/www.wiley.com\" target=\"www.wiley.com\">John Wiley &#038; Sons, Inc.<\/a>. ISSN 1532-0634. DOI <a href=\"http:\/\/dx.doi.org\/10.1002\/cpe.3805\" target=\"publication\">10.1002\/cpe.3805<\/a>. <a href=\"javascript:showAbstract('Investigating the performance of parallel applications at scale on future high-performance computing (HPC) architectures and the performance impact of different HPC architecture choices is an important component of HPC hardware\/software co-design. The Extreme-scale Simulator (xSim) is a simulation toolkit for investigating the performance of parallel applications at scale. xSim scales to millions of simulated Message Passing Interface (MPI) processes. The xSim toolkit strives to limit simulation overheads in order to maintain performance and productivity criteria. This paper documents two improvements to xSim: (1) a new deadlock resolution protocol to reduce the parallel discrete event simulation overhead, and (2) a new simulated MPI message matching algorithm to reduce the oversubscription management cost. These enhancements resulted in significant performance improvements. The simulation overhead for running the NAS Parallel Benchmark suite dropped from 1,020% to 238% for the conjugate gradient (CG) benchmark and 102% to 0% for the embarrassingly parallel (EP) benchmark. Additionally, the improvements were beneficial for reducing overheads in the highly accurate simulation mode of xSim, which is useful for resilience investigation studies for tracking intentional MPI process failures. In the highly accurate mode, the simulation overhead was reduced from 37,511% to 13,808% for CG and from 3,332% to 204% for EP.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann16new.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann16new\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Marc Snir, Robert W. Wisniewski, Jacob A. Abraham, Sarita V. Adve, Saurabh Bagchi, Pavan Balaji, Jim Belak, Pradip Bose, Franck Cappello, Bill Carlson, Andrew A. Chien, Paul Coteus, Nathan A. Debardeleben, Pedro Diniz, Christian Engelmann, Mattan Erez, Saverio Fazzari, Al Geist, Rinku Gupta, Fred Johnson, Sriram Krishnamoorthy, Sven Leyffer, Dean Liberty, Subhasish Mitra, Todd Munson, Rob Schreiber, Jon Stearley, and Eric Van Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 28, number 2, pages 127-171, May 1, 2014. <a href=\"http:\/\/www.sagepub.com\" target=\"www.sagepub.com\">SAGE Publications<\/a>. ISSN 1094-3420. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/1094342014522573\" target=\"publication\">10.1177\/1094342014522573<\/a>. <a href=\"javascript:showAbstract('We present here a report produced by a workshop on  Addressing failures in exascale computing&amp;#39; held in Park City,  Utah, 4-11 August 2012. The charter of this workshop was to  establish a common taxonomy about resilience across all the  levels in a computing system, discuss existing knowledge on  resilience across the various hardware and software layers  of an exascale system, and build on those results, examining  potential solutions from both a hardware and software  perspective and focusing on a combined approach. The workshop brought together participants with expertise in  applications, system software, and hardware; they came from  industry, government, and academia, and their interests ranged  from theory to implementation. The combination allowed broad  and comprehensive discussions and led to this document, which  summarizes and builds on those discussions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/snir14addressing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir14addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Scaling To A Million Cores And Beyond: Using Light-Weight Simulation to Understand The Challenges Ahead On The Road To Exascale<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/fgcs\" target=\"www.elsevier.com\/locate\/fgcs\">Future Generation Computer Systems (FGCS)<\/a><\/i>, volume 30, number 0, pages 59-65, January 1, 2014. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0167-739X. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.future.2013.04.014\" target=\"publication\">10.1016\/j.future.2013.04.014<\/a>. <a href=\"javascript:showAbstract('As supercomputers scale to 1,000 PFlop\/s over the next decade, investigating the performance of parallel applications at scale on future architectures and the performance impact of different architecture choices for high-performance computing (HPC) hardware\/software co-design is crucial. This paper summarizes recent efforts in designing and implementing a novel HPC hardware\/software co-design toolkit. The presented Extreme-scale Simulator (xSim) permits running an HPC application in a controlled environment with millions of concurrent execution threads while observing its performance in a simulated extreme-scale HPC system using architectural models and virtual timing. This paper demonstrates the capabilities and usefulness of the xSim performance investigation toolkit, such as its scalability to 2^27 simulated Message Passing Interface (MPI) ranks on 960 real processor cores, the capability to evaluate the performance of different MPI collective communication algorithms, and the ability to evaluate the performance of a basic Monte Carlo application with different architectural parameters.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13scaling.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann13scaling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chao Wang, Frank Mueller, Christian Engelmann, and Stephen L. Scott. <b>Proactive Process-Level Live Migration and Back Migration in HPC Environments<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 72, number 2, pages 254-267, February 1, 2012. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2011.10.009\" target=\"publication\">10.1016\/j.jpdc.2011.10.009<\/a>. <a href=\"javascript:showAbstract('As the number of nodes in high-performance computing environments keeps increasing, faults are becoming common place. Reactive fault tolerance (FT) often does not scale due to massive I\/O requirements and relies on manual job resubmission. This work complements reactive with proactive FT at the process level. Through health monitoring, a subset of node failures can be anticipated when one&amp;#39;s health deteriorates. A novel process-level live migration mechanism supports continued execution of applications during much of process migration. This scheme is integrated into an MPI execution environment to transparently sustain health-inflicted node failures, which eradicates the need to restart and requeue MPI jobs. Experiments indicate that 1-6.5 s of prior warning are required to successfully trigger live process migration while similar operating system virtualization mechanisms require 13-24 s. This self-healing approach complements reactive FT by nearly cutting the number of checkpoints in half when 70% of the faults are handled proactively. The work also provides a novel back migration approach to eliminate load imbalance or bottlenecks caused by migrated tasks. Experiments indicate the larger the amount of outstanding execution, the higher the benefit due to back migration.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang12proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#wang12proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, Christian Engelmann, and Hong H. Ong. <b>System-Level Virtualization Research at Oak Ridge National Laboratory<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/fgcs\" target=\"www.elsevier.com\/locate\/fgcs\">Future Generation Computer Systems (FGCS)<\/a><\/i>, volume 26, number 3, pages 304-307, March 1, 2010. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0167-739X. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.future.2009.07.001\" target=\"publication\">10.1016\/j.future.2009.07.001<\/a>. <a href=\"javascript:showAbstract('System-level virtualization is today enjoying a rebirth as a technique to effectively share what were then considered large computing resources to subsequently fade from the spotlight as individual workstations gained in popularity with a one machine - one user approach. One reason for this resurgence is that the simple workstation has grown in capability to rival that of anything available in the past. Thus, computing centers are again looking at the price\/performance benefit of sharing that single computing box via server consolidation. However, industry is only concentrating on the benefits of using virtualization for server consolidation (enterprise computing) whereas our interest is in leveraging virtualization to advance high-performance computing (HPC). While these two interests may appear to be orthogonal, one consolidating multiple applications and users on a single machine while the other requires all the power from many machines to be dedicated solely to its purpose, we propose that virtualization does provide attractive capabilities that may be exploited to the benefit of HPC interests. This does raise the two fundamental questions of: is the concept of virtualization (a machine sharing technology) really suitable for HPC and if so, how does one go about leveraging these virtualization capabilities for the benefit of HPC. To address these questions, this document presents ongoing studies on the usage of system-level virtualization in a HPC context. These studies include an analysis of the benefits of system-level virtualization for HPC, a presentation of research efforts based on virtualization for system availability, and a presentation of research efforts for the management of virtual systems. The basis for this document was material presented by Stephen L. Scott at the Collaborative and Grid Computing Technologies meeting held in Cancun, Mexico on April 12-14, 2007.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott10system.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott10system\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xubin (Ben) He, Li Ou, Christian Engelmann, Xin Chen, and Stephen L. Scott. <b>Symmetric Active\/Active Metadata Service for High Availability Parallel File Systems<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 69, number 12, pages 961-973, December 1, 2009. <a href=\"http:\/\/www.elsevier.com\" target=\"www.elsevier.com\">Elsevier B.V, Amsterdam, The Netherlands<\/a>. ISSN 0743-7315. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2009.08.004\" target=\"publication\">10.1016\/j.jpdc.2009.08.004<\/a>. <a href=\"javascript:showAbstract('High availability data storage systems are critical for many applications as research and business become more data-driven. Since metadata management is essential to system availability, multiple metadata services are used to improve the availability of distributed storage systems. Past research focused on the active\/standby model, where each active service has at least one redundant idle backup. However, interruption of service and even some loss of service state may occur during a fail-over depending on the used replication technique. In addition, the replication overhead for multiple metadata services can be very high. The research in this paper targets the symmetric active\/active replication model, which uses multiple redundant service nodes running in virtual synchrony. In this model, service node failures do not cause a fail-over to a backup and there is no disruption of service or loss of service state. We further discuss a fast delivery protocol to reduce the latency of the needed total order broadcast. Our prototype implementation shows that metadata service high availability can be achieved with an acceptable performance trade-off using our symmetric active\/active metadata service solution.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/he09symmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#he09symmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xubin (Ben) He, Li Ou, Martha J. Kosa, Stephen L. Scott, and Christian Engelmann. <b>A Unified Multiple-Level Cache for High Performance Cluster Storage Systems<\/b>. <i><a href=\"http:\/\/www.inderscience.com\/browse\/index.php?journalcode=ijhpcn\" target=\"www.inderscience.com\/browse\/index.php?journalcode=ijhpcn\">International Journal of High Performance Computing and Networking (IJHPCN)<\/a><\/i>, volume 5, number 1-2, pages 97-109, November 14, 2007. <a href=\"http:\/\/www.inderscience.com\" target=\"www.inderscience.com\">Inderscience Publishers, Geneve, Switzerland<\/a>. ISSN 1740-0562. DOI <a href=\"http:\/\/dx.doi.org\/10.1504\/IJHPCN.2007.015768\" target=\"publication\">10.1504\/IJHPCN.2007.015768<\/a>. <a href=\"javascript:showAbstract('Highly available data storage for high-performance computing is becoming increasingly more critical as high-end computing systems scale up in size and storage systems are developed around network-centered architectures. A promising solution is to harness the collective storage potential of individual workstations much as we harness idle CPU cycles due to the excellent price\/performance ratio and low storage usage of most commodity workstations. For such a storage system, metadata consistency is a key issue assuring storage system availability as well as data reliability. In this paper, we present a decentralized metadata management scheme that improves storage availability without sacrificing performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/he07unified.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#he07unified\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Symmetric Active\/Active High Availability for High-Performance Computing System Services<\/b>. <i><a href=\"http:\/\/www.jcomputers.us\" target=\"www.jcomputers.us\">Journal of Computers (JCP)<\/a><\/i>, volume 1, number 8, pages 43-54, December 1, 2006. <a href=\"http:\/\/www.jcomputers.us\" target=\"www.jcomputers.us\">Academy Publisher, Oulu, Finland<\/a>. ISSN 1796-203X. DOI <a href=\"http:\/\/dx.doi.org\/10.4304\/jcp.1.8.43-54\" target=\"publication\">10.4304\/jcp.1.8.43-54<\/a>. <a href=\"javascript:showAbstract('This work aims to pave the way for high availability in high-performance computing (HPC) by focusing on efficient redundancy strategies for head and service nodes. These nodes represent single points of failure and control for an entire HPC system as they render it inaccessible and unmanageable in case of a failure until repair. The presented approach introduces two distinct replication methods, internal and external, for providing symmetric active\/active high availability for multiple redundant head and service nodes running in virtual synchrony utilizing an existing process group communication system for service group membership management and reliable, totally ordered message delivery. Resented results of a prototype implementation that offers symmetric active\/active replication for HPC job and resource management using external replication show that the highest level of availability can be provided with an acceptable performance trade-off.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06symmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann06symmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, David E. Bernholdt, Narasimha R. Gottumukkala, Chokchai (Box) Leangsuksun, Jyothish Varma, Chao Wang, Frank Mueller, Aniruddha G. Shet, and Ponnuswamy (Saday) Sadayappan. <b>MOLAR: Adaptive Runtime Support for High-End Computing Operating and Runtime Systems<\/b>. <i><a href=\"http:\/\/www.sigops.org\/osr.html\" target=\"www.sigops.org\/osr.html\">ACM SIGOPS Operating Systems Review (OSR)<\/a><\/i>, volume 40, number 2, pages 63-72, April 1, 2006. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISSN 0163-5980. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1131322.1131337\" target=\"publication\">10.1145\/1131322.1131337<\/a>. <a href=\"javascript:showAbstract('MOLAR is a multi-institutional research effort that concentrates on adaptive, reliable, and efficient operating and runtime system (OS\/R) solutions for ultra-scale, high-end scientific computing on the next generation of supercomputers. This research addresses the challenges outlined in FAST-OS (forum to address scalable technology for runtime and operating systems) and HECRTF (high-end computing revitalization task force) activities by exploring the use of advanced monitoring and adaptation to improve application performance and predictability of system interruptions, and by advancing computer reliability, availability and serviceability (RAS) management systems to work cooperatively with the OS\/R to identify and preemptively resolve system issues. This paper describes recent research of the MOLAR team in advancing RAS for high-end computing OS\/Rs.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06molar.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann06molar\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Emmanuel Agullo, Mirco Altenbernd, Hartwig Anzt, Leonardo Bautista-Gomez, Tommaso Benacchio, Luca Bonaventura, Hans-Joachim Bungartz, Sanjay Chatterjee, Florina M. Ciorba, Nathan DeBardeleben, Daniel Drzisga, Sebastian Eibl, Christian Engelmann, Wilfried N. Gansterer, Luc Giraud, Dominik G&ouml;ddeke, Marco Heisig, Fabienne J&eacute;z&eacute;quel, Nils Kohl, Xiaoye Sherry Li, Romain Lion, Miriam Mehl, Paul Mycek, Michael Obersteiner, Enrique S. Quintana-Ort&iacute;, Francesco&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":16,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-18","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/18","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=18"}],"version-history":[{"count":10,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/18\/revisions"}],"predecessor-version":[{"id":1236,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/18\/revisions\/1236"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=18"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}