{"id":96,"date":"2026-09-04T08:00:54","date_gmt":"2026-09-04T08:00:54","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=96"},"modified":"2026-09-05T01:06:51","modified_gmt":"2026-09-05T01:06:51","slug":"peer-reviewed-workshop-papers","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=96","title":{"rendered":"Peer-Reviewed Workshop Papers"},"content":{"rendered":"<ol>\n<li>Christian Engelmann, Andrew Ayres, Stephen DeWitt, Michael J. Brim, and Brett Eiffert. <b>Building Resilient Self-Driving Laboratories with the INTERSECT Federated Ecosystem<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc26.supercomputing.org\" target=\"sc26.supercomputing.org\">39th International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2026<\/a>: <a href=\"http:\/\/wordpress.cels.anl.gov\/xloop-2026\/\" target=\"wordpress.cels.anl.gov\/xloop-2026\/\">8th Annual Workshop on Extreme-Scale Experiment-in-the-Loop Computing (XLOOP) 2026<\/a><\/i>, Chicago, IL, USA, November 15, 2026. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. To appear. <a href=\"javascript:showAbstract('Failure resilience in federated ecosystems for instrument science presents a critical challenge. Failures disrupt experiments and make them potentially useless, wasting valuable resources and creating setbacks. Oak Ridge National Laboratory&amp;#39;s Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) offers a federated ecosystem for instrument science, enabling autonomous experiments, self-driving laboratories, smart manufacturing, and AI-driven design, discovery, and evaluation. This paper documents the recent advances in creating a resilient INTERSECT ecosystem. The proposed solution includes a resilient architecture with resilience design patterns, a resilient system of systems (SoS) architecture, and a resilient microservices architecture; and a resilient software development kit with reliable service communication and asynchronous and synchronous failure detection and notification. The resilience capabilities are demonstrated for an autonomous additive manufacturing process with a real-time feedback loop.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#engelmann26building\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen Boehm, Craig A. Bridges, Patrick Widener, Terry Jones, Sheikh Ghafoor, Christian Engelmann, and Olga Kuchar. <b>The INTERSECT Scientific Data Layer: An Ontological Framework for Data Provenance for Complex Scientific Workflows<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/2026.euro-par.org\" target=\"2026.euro-par.org\">32nd European Conference on Parallel and Distributed Computing (Euro-Par) 2026 Workshops<\/a>: <a href=\"http:\/\/www.hipes-workshop.org\/\" target=\"www.hipes-workshop.org\/\">3rd Workshop on High-Performance eScience Tools and Applications (HiPES)<\/a><\/i>, Pisa, Italy, August 25, 2026. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. To appear. <a href=\"javascript:showAbstract('Complex scientific workflows have multiple stages, including experiments, simulations, data analyses, and visualization, generating data and metadata stored across heterogeneous storage infrastructures and used in downstream stages or future experimental campaigns. This paper presents the design and implementation of a comprehensive ontological framework for managing scientific data generated within such workflows, addressing the key challenges of data interoperability, provenance capture, and adherence to Findable, Accessible, Interoperable, and Reusable (FAIR) data principles. The Autonomous Chemistry Laboratory (ACL) at Oak Ridge National Laboratory (ORNL) enables automated liquid phase and solid state synthesis and related chemical analysis. Our framework has been deployed within the ACL for a native and machine-interpretable semantic representation of its ecosystem, including instrument capabilities, synthesis workflows, analytical observations, and experimental results. The proposed approach enables end-to-end provenance tracking, supports heterogeneous data formats, and establishes a foundation for Artificial Intelligence (AI)-ready scientific discovery. We validate the framework through concrete modeling examples.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#boehm28intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Olivera Kotevska, Trong Nguyen, Rafael Ferreira da Silva, Christian Engelmann, and Prasanna Balaprakash. <b>Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/2026.eurosys.org\/\" target=\"2026.eurosys.org\/\">21st European Conference on Computer Systems (EuroSyS)<\/a>: <a href=\"http:\/\/euromlsys.eu\/\" target=\"euromlsys.eu\/\">6th European Workshop on Machine Learning and Systems (EuroMLSys)<\/a><\/i>, pages 439-446, Edinburgh, United Kingdom, April 27, 2026. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 979-8-4007-2605-7. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3805621.3807639\" target=\"publication\">10.1145\/3805621.3807639<\/a>. Acceptance rate 69.2% (18\/26). <a href=\"javascript:showAbstract('Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronization and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kotevska26scalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kotevska26scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Michael J. Brim, Lance Drane, Marshall McDonnell, Christian Engelmann, and Addi Malviya Thakur. <b>A Microservices Architecture Toolkit for Interconnected Science Ecosystems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc24.supercomputing.org\" target=\"sc24.supercomputing.org\">37th International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2024<\/a>: <a href=\"http:\/\/works-workshop.org\/\" target=\"works-workshop.org\/\">19th Workshop on Workflows in Support of Large-Scale  Science (WORKS) 2024<\/a><\/i>, pages 2072-2079, Atlanta, GA, USA, November 18, 2024. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 979-8-3503-5554-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SCW63240.2024.00259\" target=\"publication\">10.1109\/SCW63240.2024.00259<\/a>. Acceptance rate 66.7% (10\/15). <a href=\"javascript:showAbstract('Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for integrating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architecture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability development, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/brim24microservices.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#brim24microservices\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar and Christian Engelmann. <b>RDPM: An Extensible Tool for Resilience Design Patterns Modeling<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/2021.euro-par.org\" target=\"2021.euro-par.org\">27th European Conference on Parallel and Distributed Computing (Euro-Par) 2021 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2021\">14th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 283-297, Lisbon, Portugal, August 30, 2021. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-031-06155-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-031-06156-1_23\" target=\"publication\">10.1007\/978-3-031-06156-1_23<\/a>. Acceptance rate 66.7% (4\/6). <a href=\"javascript:showAbstract('Resilience to faults, errors, and failures in extreme-scale HPC systems is a critical challenge. Resilience design patterns offer a new, structured hardware and software design approach for improving resilience. While prior work focused on developing performance, reliability, and availability models for resilience design patterns, this paper extends it by providing a Resilience Design Patterns Modeling (RDPM) tool which allows (1) exploring performance, reliability, and availability of each resilience design pattern, (2) offering customization of parameters to optimize performance, reliability, and availability, and (3) allowing investigation of trade-off models for combining multiple patterns for practical resilience solutions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar21rdpm.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar21rdpm\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mohit Kumar and Christian Engelmann. <b>Models for Resilience Design Patterns<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc20.supercomputing.org\" target=\"sc20.supercomputing.org\">33rd International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2020<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2020\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2020\">10th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2020<\/a><\/i>, pages 21-30, Atlanta, GA, USA, November 11, 2020. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7381-1080-6. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS51974.2020.00008\" target=\"publication\">10.1109\/FTXS51974.2020.00008<\/a>. Acceptance rate 66.7% (6\/9). <a href=\"javascript:showAbstract('Resilience plays an important role in supercomputers by providing correct and efficient operation in case of faults, errors, and failures. Resilience design patterns offer blueprints for effectively applying resilience technologies. Prior work focused on developing initial efficiency and performance models for resilience design patterns. This paper extends it by (1) describing performance, reliability, and availability models for all structural resilience design patterns, (2) providing more detailed models that include flowcharts and state diagrams, and (3) introducing the Resilience Design Pattern Modeling (RDPM) tool that calculates and plots the performance, reliability, and availability metrics of individual patterns and pattern combinations.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar20models.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/kumar20models.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#kumar20models\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Piyush Sao, Christian Engelmann, Srinivas Eswar, Oded Green, and Richard Vuduc. <b>Self-stabilizing Connected Components<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc19.supercomputing.org\" target=\"sc19.supercomputing.org\">32nd International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2019<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2019\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2019\">9th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2019<\/a><\/i>, pages 50-59, Denver, CO, USA, November 22, 2019. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-6013-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS49593.2019.00011\" target=\"publication\">10.1109\/FTXS49593.2019.00011<\/a>. Acceptance rate 60.0% (6\/10). <a href=\"javascript:showAbstract('For the problem of computing the connected components of a graph, this paper considers the design of algorithms that are resilient to transient hardware faults, like bit flips. More specifically, it applies the technique of \\emphself-stabilization. A system is self-stabilizing if, when starting from a valid or invalid state, it is guaranteed to reach a valid state after a finite number of steps. Therefore on a machine subject to a transient fault, a self-stabilizing algorithm could recover if that fault caused the system to enter an invalid state. We give a comprehensive analysis of the valid and invalid states during label propagation and derive algorithms to verify and correct the invalid state. The self-stabilizing label-propagation algorithm performs &amp;#36;\\bigoV &amp;#322;og V additional computation and requires \\bigoV additional storage over its conventional counterpart (and, as such, does not increase asymptotic complexity over conventional). When run against a battery of simulated fault injection tests, the self-stabilizing label propagation algorithm exhibits more resilient behavior than a triple modular redundancy (TMR) based fault-tolerant algorithm in 80% of cases. From a performance perspective, it also outperforms TMR as it requires fewer iterations in total. Beyond the fault-tolerance properties of self-stabilizing label-propagation, we believe, they are useful from the theoretical perspective; and may have other use-cases.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/sao19self-stabilizing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/sao19self-stabilizing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#sao19self-stabilizing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Geoffroy R. Vall&eacute;e, and Swaroop Pophale. <b>Concepts for OpenMP Target Offload Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/parallel.auckland.ac.nz\/iwomp2019\" target=\"parallel.auckland.ac.nz\/iwomp2019\">15th International Workshop on OpenMP (IWOMP) 2019<\/a><\/i>, pages 78-93, Auckland, New Zealand, September 11-13, 2019. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-030-28595-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-28596-8_6\" target=\"publication\">10.1007\/978-3-030-28596-8_6<\/a>. <a href=\"javascript:showAbstract('Recent reliability issues with one of the fastest supercomputers in the world, Titan at Oak Ridge National Laboratory, demonstrated the need for resilience in large-scale heterogeneous computing. OpenMP currently does not address error and failure behavior. This paper takes a first step toward resilience for heterogeneous systems by providing the concepts for resilient OpenMP offload to devices. Using real-world error and failure observations, the paper describes the concepts and terminology for resilient OpenMP target offload, including error and failure classes and resilience strategies. It details the experienced general-purpose computing on graphics processing units errors and failures in Titan. It further proposes improvements in OpenMP, including a preliminary prototype design, to support resilient offload to devices for efficient handling of errors and failures in heterogeneous high-performance computing systems');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann19concepts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann19concepts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann19concepts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>A Comprehensive Informative Metric for Analyzing HPC System Status using the LogSCAN Platform<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc18.supercomputing.org\" target=\"sc18.supercomputing.org\">31st International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\">8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2018<\/a><\/i>, pages 29-38, Dallas, TX, USA, November 16, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-0222-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS.2018.00007\" target=\"publication\">10.1109\/FTXS.2018.00007<\/a>. Acceptance rate 45.0% (9\/20). <a href=\"javascript:showAbstract('Log processing by Spark and Cassandra-based ANalytics (LogSCAN) is a newly developed analytical platform that provides flexible and scalable data gathering, transformation and computation. One major challenge is to effectively summarize the status of a complex computer system, such as the Titan supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). Although there is plenty of operational and maintenance information collected and stored in real time, which may yield insights about short- and long-term system status, it is difficult to present this information in a comprehensive form. In this work, we present system information entropy (SIE), a newly developed metric that leverages the powers of traditional machine learning techniques and information theory. By compressing the multi-variant multi-dimensional event information recorded during the operation of the targeted system into a single time series of SIE, we demonstrate that the historical system status can be sensitively represented concisely and comprehensively. Given a sharp indicator as SIE, we argue that follow-up analytics based on SIE will reveal in-depth knowledge about system status using other sophisticated approaches, such as pattern recognition in the temporal domain or causality analysis incorporating extra independent metrics of the system.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18comprehensive2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hui18comprehensive2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hui18comprehensive2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf and Christian Engelmann. <b>Analyzing the Impact of System Reliability Events on Applications in the Titan Supercomputer<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc18.supercomputing.org\" target=\"sc18.supercomputing.org\">31st International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2018\">8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2018<\/a><\/i>, pages 39-48, Dallas, TX, USA, November 16, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-7281-0222-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/FTXS.2018.00008\" target=\"publication\">10.1109\/FTXS.2018.00008<\/a>. Acceptance rate 45.0% (9\/20). <a href=\"javascript:showAbstract('Extreme-scale computing systems employ Reliability, Availability and Serviceability (RAS) mechanisms and infrastructure to log events from multiple system components. In this paper, we analyze RAS logs in conjunction with the  application placement and scheduling database, in order to  understand the impact of common RAS events on application performance. This study conducted on the records of about 2 million applications executed on Titan supercomputer provides important insights for system users, operators and computer science researchers. In this paper, we investigate the impact of RAS events on application performance and its variability by comparing cases where events are recorded with corresponding cases where no events are recorded. Such a statistical investigation is possible since we observed that system users tend to execute their applications multiple times. Our analysis reveals that most RAS events do impact application performance, although not always. We also find that different system components affect application performance differently. In particular, our investigation includes the following components: parallel file system, processor, memory, graphics processing units, system and user software issues. Our work establishes the importance of providing feedback to system users for increasing operational efficiency of extreme-scale systems.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18analyzing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18analyzing.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18analyzing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Byung Hoon (Hoony) Park, Yawei Hui, Swen Boehm, Rizwan Ashraf, Christian Engelmann, and Christopher Layton. <b>A Big Data Analytics Framework for HPC Log Data: Three Case Studies Using the Titan Supercomputer Log<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster2018.github.io\" target=\"cluster2018.github.io\">19th IEEE International Conference on Cluster Computing (Cluster) 2018<\/a>: <a href=\"http:\/\/sites.google.com\/site\/hpcmaspa2018\" target=\"sites.google.com\/site\/hpcmaspa2018\">5th Workshop on Monitoring and Analysis for High Performance Systems Plus Applications (HPCMASPA) 2018<\/a><\/i>, pages 571-579, Belfast, UK, September 10, 2018. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-8319-4. ISSN 2168-9253. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTER.2018.00073\" target=\"publication\">10.1109\/CLUSTER.2018.00073<\/a>. <a href=\"javascript:showAbstract('Reliability, availability and serviceability (RAS) logs of high performance computing (HPC) resources, when closely investigated in spatial and temporal dimensions, can provide invaluable information regarding system status, performance, and resource utilization. These data are often generated from multiple logging systems and sensors that cover many components of the system. The analysis of these data for finding persistent temporal and spatial insights faces two main difficulties: the volume of RAS logs makes manual inspection difficult and the unstructured nature and unique properties of log data produced by each subsystem adds another dimension of difficulty in identifying implicit correlation among recorded events. To address these issues, we recently developed a multi-user Big Data analytics framework for HPC log data at Oak Ridge National Laboratory (ORNL). This paper introduces three in-progress data analytics projects that leverage this framework to assess system status, mine event patterns, and study correlations between user applications and system events. We describe the motivation of each project and detail their workflows using three years of log data collected from ORNL&amp;#39;s Titan supercomputer.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/park18big.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/park18big.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#park18big\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Rizwan Ashraf and Christian Engelmann. <b>Performance Efficient Multiresilience using Checkpoint Recovery in Iterative Algorithms<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2018.org\" target=\"europar2018.org\">24th European Conference on Parallel and Distributed Computing (Euro-Par) 2018 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2018\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2018\">11th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 813-825, Turin, Italy, August 28, 2018. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-030-10549-5. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-10549-5_63\" target=\"publication\">10.1007\/978-3-030-10549-5_63<\/a>. Acceptance rate 50.0% (4\/8). <a href=\"javascript:showAbstract('In this paper, we address the design challenge of building multiresilient iterative high-performance computing (HPC) applications. Multiresilience in HPC applications is the ability to tolerate and maintain forward progress in the presence of both soft errors and process failures. We address the challenge by proposing performance models which are useful to design performance efficient and resilient iterative applications. The models consider the interaction between soft error and process failure resilience solutions. We experimented with a linear solver application with two distinct kinds of soft error detectors: one detector is high overhead and high accuracy, whereas the second is low overhead and low accuracy. We show how both can be leveraged for verifying the integrity of checkpointed state used to recover from both soft errors and process failures. Our results show the performance efficiency and resiliency benefit of employing the low overhead detector with high frequency within the checkpoint interval, so that timely soft error recovery can take place, resulting in less re-computed work.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ashraf18performance.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ashraf18performance.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ashraf18performance\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Byung Hoon (Hoony) Park, Saurabh Hukerikar, Christian Engelmann, and Ryan Adamson. <b>Big Data Meets HPC Log Analytics: Scalable Approach to Understanding Systems at Extreme Scale<\/b>. In <i>Proceedings of the <a href=\"http:\/\/cluster17.github.io\" target=\"cluster17.github.io\">18th IEEE International Conference on Cluster Computing (Cluster) 2017<\/a>: <a href=\"http:\/\/sites.google.com\/site\/hpcmaspa2017\" target=\"sites.google.com\/site\/hpcmaspa2017\">4th Workshop on Monitoring and Analysis for High Performance Systems Plus Applications (HPCMASPA) 2017<\/a><\/i>, pages 758-765, Honolulu, HI, USA, September 5, 2017. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-5386-2327-5. ISSN 2168-9253. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTER.2017.113\" target=\"publication\">10.1109\/CLUSTER.2017.113<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s high-performance computing (HPC) systems are heavily instrumented generating logs containing information about abnormal events, such as critical conditions, faults, errors and failures, system resource utilization, and about the resource usage of user applications. These logs, once fully analyzed and correlated, can produce detailed information about the system health, root causes of failures, and analyze an application's interactions with the system, providing invaluable insights to domain scientists and system administrators. However, processing HPC logs requires deep understanding of hardware and software components at multiple layers of the system stack. Moreover, most log data is unstructured and voluminous, making it more difficult for scientists and engineers to analyze the data. With rapid increases in the scale and complexity of HPC systems, log data processing is becoming a big data challenge. This paper introduces a HPC log data analytics framework that is based on a distributed NoSQL database technology, which provides scalability and high availability, and the Apache Spark for rapid in-memory processing of log data. The framework enables the extraction of a range of information about the system so that system administrators and end users alike can obtain necessary insights for their specific needs. We describe our experience with using this framework to glean insights from the log data derived from the Titan supercomputer at the Oak Ridge National Laboratory.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/park17big.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/park17big.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#park17big\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Pattern-based Modeling of High-Performance Computing Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2017.usc.es\" target=\"europar2017.usc.es\">23rd European Conference on Parallel and Distributed Computing (Euro-Par) 2017 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2017\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2017\">10th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 557-568, Santiago de Compostela, Spain, August 29, 2017. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-75177-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-75178-8_45\" target=\"publication\">10.1007\/978-3-319-75178-8_45<\/a>. Acceptance rate 66.7% (4\/6). <a href=\"javascript:showAbstract('The design of supercomputing systems and their applications must consider resilience, and power consumption as the key design parameters when designing to achieve higher performance. In previous work, we established a structured methodology for developing resilience solutions based on the concept of design patterns. In this paper we discuss analytical models for the design patterns to support quantitative analysis of their performance and reliability characteristics.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17pattern-based.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar17pattern-based.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar17pattern-based\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar, Rizwan Ashraf, and Christian Engelmann. <b>Towards New Metrics for High-Performance Computing Resilience<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.hpdc.org\/2017\" target=\"www.hpdc.org\/2017\">26th ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC) 2017<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2017\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2017\">7th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2017<\/a><\/i>, pages 23-30, Washington, D.C., June 26-30, 2017. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-5001-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3086157.3086163\" target=\"publication\">10.1145\/3086157.3086163<\/a>. Acceptance rate 83.3% (5\/6). <a href=\"javascript:showAbstract('Ensuring the reliability of applications is becoming an increasingly important challenge as high-performance computing (HPC) systems experience an ever-growing number of faults, errors and failures. While the HPC community has made substantial progress in developing various resilience solutions, it continues to rely on platform-based metrics to quantify application resiliency improvements. The resilience of an HPC application is concerned with the reliability of the application outcome as well as the fault handling efficiency. To understand the scope of impact, effective coverage and performance efficiency of existing and emerging resilience solutions, there is a need for new metrics. In this paper, we develop new ways to quantify resilience that consider both the reliability and the performance characteristics of the solutions from the perspective of HPC applications. As HPC systems continue to evolve in terms of scale and complexity, it is expected that applications will experience various types of faults, errors and failures, which will require applications to apply multiple resilience solutions across the system stack. The proposed metrics are intended to be useful for understanding the combined impact of these solutions on an application&amp;#39;s ability to produce correct results and to evaluate their overall impact on an application's performance in the presence of various modes of faults.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar17towards.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar17towards.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar17towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Saurabh Hukerikar and Christian Engelmann. <b>Language Support for Reliable Memory Regions<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/lcpc2016.wordpress.com\" target=\"lcpc2016.wordpress.com\">29th International Workshop on Languages and Compilers for Parallel Computing<\/a><\/i>, pages 73-87, Rochester, NY, USA, September 28-30, 2016. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-52708-6. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-52709-3_6\" target=\"publication\">10.1007\/978-3-319-52709-3_6<\/a>. Acceptance rate 76.9% (20\/26). <a href=\"javascript:showAbstract('The path to exascale computational capabilities in high-performance computing (HPC) systems is challenged by the evolution of the architectures of supercomputing systems. The constraints of power have driven designs that include increasingly heterogeneous architectures and complex memory hierarchies. These systems are also expected to experience in an increased rate of errors, such that the applications will no longer be able to assume correct behavior of the underlying machine. To enable the scientific community to succeed in scaling their applications and harness the capabilities of exascale systems, we need software strategies that provide mechanisms for explicit management of locality and resilience to errors in the system. In prior work, we introduced the concept of explicitly reliable memory regions, called havens. Memory management using havens supports selective reliability through a region-based approach to memory allocation. Havens enable the creation of explicit software-enabled robust memory containers for which resilient behavior is guaranteed. In this paper, we propose language support for havens through type annotations that make the structure of a program&amp;#39;s havens more explicit. We describe how the extended haven-based memory management model is implemented and the impact on the resiliency of a conjugate gradient application.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hukerikar16language.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/hukerikar16language.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#hukerikar16language\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Christian Engelmann, Geoffroy Vall&eacute;e, Ferrol Aderholdt, and Stephen L. Scott. <b>A Cooperative Approach to Virtual Machine Based Fault Injection<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2016.inria.fr\" target=\"europar2016.inria.fr\">22nd European Conference on Parallel and Distributed Computing (Euro-Par) 2016 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2016\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2016\">9th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 671-682, Grenoble, France, August 23, 2016. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-58943-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-58943-5_54\" target=\"publication\">10.1007\/978-3-319-58943-5_54<\/a>. Acceptance rate 55.6% (5\/9). <a href=\"javascript:showAbstract('Resilience investigations often employ fault injection (FI) tools to study the effects of simulated errors on a target system. It is important to keep the target system under test (SUT) isolated from the controlling environment in order to maintain control of the experiment. Virtual machines (VMs) have been used to aid these investigations due to the strong isolation properties of system-level virtualization. A key challenge in fault injection tools is to gain proper insight and context about the SUT. In VM-based FI tools, this challenge of target con- text is increased due to the separation between host and guest (VM). We discuss an approach to VM-based FI that leverages virtual machine introspection (VMI) methods to gain insight into the target&amp;#39;s context running within the VM. The key to this environment is the ability to provide basic information to the FI system that can be used to create a map of the target environment. We describe a proof- of-concept implementation and a demonstration of its use to introduce simulated soft errors into an iterative solver benchmark running in user-space of a guest VM.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton16cooperative.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton16cooperative.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton16cooperative\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Zachary Parchman, Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, and David E. Bernholdt. <b>Adding Fault Tolerance to NPB Benchmarks Using ULFM<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.hpdc.org\/2016\" target=\"www.hpdc.org\/2016\">25th ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC) 2016<\/a>: <a href=\"http:\/\/sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2016\" target=\"sites.google.com\/site\/ftxsworkshop\/home\/ftxs-2016\">6th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS) 2016<\/a><\/i>, pages 19-26, Kyoto, Japan, May 31 &#8211; June 4, 2016. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-4503-4349-7. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/2909428.2909429\" target=\"publication\">10.1145\/2909428.2909429<\/a>. Acceptance rate 85.7% (6\/7). <a href=\"javascript:showAbstract('In the world of high-performance computing, fault tolerance and application resilience are becoming some of the primary concerns because of increasing hardware failures and memory corruptions. While the research community has been investigating various options, from system-level solutions to application-level solutions, standards such as the Message Passing Interface (MPI) are also starting at including such capabilities. The current proposal for MPI fault tolerant is centered around the User-Level Failure Mitigation (ULFM) concept, which provides means for fault detection and recovery of the MPI layer. This approach does not address application-level recovery, which is current left to application developers. In this work, we present a modification of some of the benchmarks of the NAS parallel benchmark (NPB) to include support of the ULFM capabilities as well as application- level strategies and mechanisms for application-level failure recovery. As such, we present: (i) an application-level library to &amp;#34;checkpoint&amp;#34; data, (ii) extensions of NPB benchmarks for fault tolerance based on different strategies, (iii) a fault injection tool, and (iv) some preliminary experiments that shows the impact of such fault tolerant strategies on the application execution.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/parchman16adding.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/parchman16adding.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#parchman16adding\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Garry Smith, Christian Engelmann, Geoffroy Vall&eacute;e, Ferrol Aderholdt, and Stephen L. Scott. <b>What is the right balance for performance and isolation with virtualization in HPC?<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2014.dcc.fc.up.pt\" target=\"europar2014.dcc.fc.up.pt\">20th European Conference on Parallel and Distributed Computing (Euro-Par) 2014 Workshops<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/Resilience\/2014\" target=\"www.csm.ornl.gov\/srt\/conferences\/Resilience\/2014\">7th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 570-581, Porto, Portugal, August 25, 2014. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-14325-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-14325-5_49\" target=\"publication\">10.1007\/978-3-319-14325-5_49<\/a>. Acceptance rate 60.0% (6\/10). <a href=\"javascript:showAbstract('The use of virtualization in high-performance computing (HPC) has been suggested as a means to provide tailored services and added functionality that many users expect from full-featured Linux cluster environments. While the use of virtual machines in HPC can offer several benefits, maintaining performance is a crucial factor. In some instances performance criteria are placed above isolation properties and selective relaxation of isolation for performance is an important characteristic when considering resilience for HPC environments employing virtualization. In this paper we consider some of the factors associated with balancing performance and isolation in configurations that employ virtual machines. In this context, we propose a classification of errors based on the concept of &amp;#34;error zones&amp;#34;, as well as a detailed analysis of the trade-offs between resilience and performance based on the level of isolation provided by virtualization solutions. Finally, the results from a set of experiments are presented, that use different virtualization solutions, and in doing so allow further elucidation of the topic.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton14what.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton14what.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton14what\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>Toward a Performance\/Resilience Tool for Hardware\/Software Co-Design of High-Performance Computing Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icpp2013.ens-lyon.fr\" target=\"icpp2013.ens-lyon.fr\">42nd International Conference on Parallel Processing (ICPP) 2013<\/a>: <a href=\"http:\/\/www.psti-workshop.org\" target=\"www.psti-workshop.org\">4th International Workshop on Parallel Software Tools and Tool Infrastructures (PSTI)<\/a><\/i>, pages 962-971, Lyon, France, October 2, 2013. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-5117-3. ISSN 0190-3918. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICPP.2013.114\" target=\"publication\">10.1109\/ICPP.2013.114<\/a>. <a href=\"javascript:showAbstract('xSim is a simulation-based performance investigation toolkit that permits running high-performance computing (HPC) applications in a controlled environment with millions of concurrent execution threads, while observing application performance in a simulated extreme-scale system for hardware\/software co-design. The presented work details newly developed features for xSim that permit the injection of MPI process failures, the propagation\/detection\/notification of such failures within the simulation, and their handling using application-level checkpoint\/restart. These new capabilities enable the observation of application behavior and performance under failure within a simulated future-generation HPC system using the most common fault handling technique.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13toward.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann13toward.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann13toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mahesh Lagadapati, Frank Mueller, and Christian Engelmann. <b>Tools for Simulation and Benchmark Generation at Exascale<\/b>. In <i>Proceedings of the <a href=\"http:\/\/tools.zih.tu-dresden.de\/2013\/\" target=\"tools.zih.tu-dresden.de\/2013\/\">7th Parallel Tools Workshop<\/a><\/i>, pages 19-24, Dresden, Germany, September 3-4, 2013. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-319-08143-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-319-08144-1_2\" target=\"publication\">10.1007\/978-3-319-08144-1_2<\/a>. <a href=\"javascript:showAbstract('The path to exascale high-performance computing (HPC) poses several challenges related to power, performance, resilience, productivity, programmability, data movement, and data management. Investigating the performance of parallel applications at scale on future architectures and the performance impact of different architecture choices is an important component of HPC hardware\/software co-design. Simulations using models of future HPC systems and communication traces from applications running on existing HPC systems can offer an insight into the performance of future architectures. This work targets technology developed for scalable application tracing of communication events and memory profiles, but can be extended to other areas, such as I\/O, control flow, and data flow. It further focuses on extreme-scale simulation of millions of Message Passing Interface (MPI) ranks using a lightweight parallel discrete event simulation (PDES) toolkit for performance evaluation. Instead of simply replaying a trace within a simulation, the approach is to generate a benchmark from it and to run this benchmark within a simulation using models to reflect the performance characteristics of future-generation HPC systems. This provides a number of benefits, such as eliminating the data intensive trace replay and enabling simulations at different scales. The presented work utilizes the ScalaTrace tool to generate scalable trace files, the ScalaBenchGen tool to generate the benchmark, and the xSim tool to run the benchmark within a simulation.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/lagadapati13tools.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/lagadapati13tools.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#lagadapati13tools\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Swen B&ouml;hm, Christian Engelmann, and Geoffroy Vall&eacute;e. <b>Using Performance Tools to Support Experiments in HPC Resilience<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.europar2013.org\/\" target=\"www.europar2013.org\/\">19th European Conference on Parallel and Distributed Computing (Euro-Par) 2013 Workshops<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2013\" target=\"xcr.cenit.latech.edu\/resilience2013\">6th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 727-736, Aachen, Germany, August 26, 2013. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-54419-4. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-54420-0_71\" target=\"publication\">10.1007\/978-3-642-54420-0_71<\/a>. Acceptance rate 87.5% (7\/8). <a href=\"javascript:showAbstract('The high performance computing (HPC) community is working to address fault tolerance and resilience concerns for current and future large scale computing platforms. This is driving enhancements in the programming environments, specifically research on enhancing message passing libraries to support fault tolerant computing capabilities. The community has also recognized that tools for resilience experimentation are greatly lacking. However, we argue that there are several parallels between &amp;#34;performance tools&amp;#34; and ``resilience tools&amp;#39;'. As such, we believe the rich set of HPC performance-focused tools can be extended (repurposed) to benefit the resilience community. In this paper, we describe the initial motivation to leverage standard HPC performance analysis techniques to aid in developing diagnostic tools to assist fault tolerance experiments for HPC applications. These diagnosis procedures help to provide context for the system when the errors (failures) occurred. We describe our initial work in leveraging an MPI performance trace tool to assist in providing global context during fault injection experiments. Such tools will assist the HPC resilience community as they extend existing and new application codes to support fault tolerances.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton13using.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton13using.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton13using\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Ian S. Jones and Christian Engelmann. <b>Simulation of Large-Scale HPC Architectures<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icpp2011.org\" target=\"icpp2011.org\">40th International Conference on Parallel Processing (ICPP) 2011<\/a>: <a href=\"http:\/\/www.psti-workshop.org\" target=\"www.psti-workshop.org\">2nd International Workshop on Parallel Software Tools and Tool Infrastructures (PSTI)<\/a><\/i>, pages 447-456, Taipei, Taiwan, September 13-19, 2011. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-4511-0. ISSN 1530-2016. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICPPW.2011.44\" target=\"publication\">10.1109\/ICPPW.2011.44<\/a>. <a href=\"javascript:showAbstract('The Extreme-scale Simulator (xSim) is a recently developed performance investigation toolkit that permits running high-performance computing (HPC) applications in a controlled environment with millions of concurrent execution threads. It allows observing parallel application performance properties in a simulated extreme-scale HPC system to further assist in HPC hardware and application software co-design on the road toward multi-petascale and exascale computing. This paper presents a newly implemented network model for the xSim performance investigation toolkit that is capable of providing simulation support for a variety of HPC network architectures with the appropriate trade-off between simulation scalability and accuracy. The taken approach focuses on a scalable distributed solution with latency and bandwidth restrictions for the simulated network. Different network architectures, such as star, ring, mesh, torus, twisted torus and tree, as well as hierarchical combinations, such as to simulate network-on-chip and network-on-node, are supported. Network traffic congestion modeling is omitted to gain simulation scalability by reducing simulation accuracy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/jones11simulation.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/jones11simulation.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#jones11simulation\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Kurt Ferreira, Frank Mueller, and Christian Engelmann. <b>A Tunable, Software-based DRAM Error Detection and Correction Library for HPC<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2011.bordeaux.inria.fr\/\" target=\"europar2011.bordeaux.inria.fr\/\">17th European Conference on Parallel and Distributed Computing (Euro-Par) 2011 Workshops, Part II<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2011\" target=\"xcr.cenit.latech.edu\/resilience2011\">4th Workshop on Resiliency in High Performance Computing (Resilience) in Clusters, Clouds, and Grids<\/a><\/i>, pages 251-261, Bordeaux, France, August 29 &#8211; September 2, 2011. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-29740-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-29740-3_29\" target=\"publication\">10.1007\/978-3-642-29740-3_29<\/a>. Acceptance rate 60.0% (12\/20). <a href=\"javascript:showAbstract('Proposed exascale systems will present a number of considerable resiliency challenges. In particular, DRAM soft-errors, or bit-flips, are expected to greatly increase due to the increased memory density of these systems. Current hardware-based fault-tolerance methods will be unsuitable for addressing the expected soft error frequency rate. As a result, additional software will be needed to address this challenge. In this paper we introduce LIBSDC, a tunable, transparent silent data corruption detection and correction library for HPC applications. LIBSDC provides comprehensive SDC protection for program memory by implementing on-demand page integrity verification. Experimental benchmarks with Mantevo HPCCG show that once tuned, LIBSDC is able to achieve SDC protection with 50% overhead of resources, less than the 100% needed for double modular redundancy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala11tunable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#fiala11tunable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Geoffroy R. Vall&eacute;e, Christian Engelmann, and Stephen L. Scott. <b>A Case for Virtual Machine based Fault Injection in a High-Performance Computing Environment<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2011.bordeaux.inria.fr\/\" target=\"europar2011.bordeaux.inria.fr\/\">17th European Conference on Parallel and Distributed Computing (Euro-Par) 2011<\/a>: <a href=\"http:\/\/www.csm.ornl.gov\/srt\/conferences\/hpcvirt2011\" target=\"www.csm.ornl.gov\/srt\/conferences\/hpcvirt2011\">5th Workshop on System-level Virtualization for High Performance Computing (HPCVirt)<\/a><\/i>, pages 234-243, Bordeaux, France, August 29 &#8211; September 2, 2011. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-29737. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-29737-3_27\" target=\"publication\">10.1007\/978-3-642-29737-3_27<\/a>. <a href=\"javascript:showAbstract('Large-scale computing platforms provide tremendous capabilities for scientific discovery. These systems have hundreds of thousands of computing cores, hundreds of terabytes of memory, and enormous high-performance interconnection networks. These systems are facing enormous challenges to achieve performance at such scale. Failures are an Achilles heel of these enormous systems. As applications and system software scale up to multi-petaflop and beyond to exascale platforms, the occurrence of failure will be much more common. This has given rise to a push in fault-tolerance and resilience research for HPC systems. This includes work on log analysis to identify types of failures, enhancements to the Message Passing Interface (MPI) to incorporate fault awareness, and a variety of fault tolerance mechanisms that span redundant computation, algorithm based fault tolerance, and advanced checkpoint\/ restart techniques. While there is much work to be done on the FT\/Resilience mechanisms for such large-scale systems, there is also a profound gap in the tools for experimentation. This gap is compounded by the fact that HPC environments have stringent performance requirements and are often highly customized. The tool chain for these systems are often tailored for the platform and while the majority of systems on the Top500 Supercomputer list run Linux, these operating environments typically contain many site\/machine specific enhancements. Therefore, it is desirable to maintain a consistent execution environment to minimize end-user (scientist) interruption. The work on system-level virtualization for HPC system offers a unique opportunity to maintain a consistent execution environment via a virtual machine (VM). Recent work on virtualization for HPC has shown that low-overhead, high performance systems can be realized [1, 2] Virtualization also provides a clean abstraction for building experimental tools for investigation into the effects of failures in HPC and the related research on FT\/ Resilience mechanisms and policies. In this paper we discuss the motivation for tools to perform fault injection in an HPC context, and outline an approach that can leverage virtualization.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton11case.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton11case.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton11case\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Frank Lauer. <b>Facilitating Co-Design for Extreme-Scale Systems Through Lightweight Simulation<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.cluster2010.org\" target=\"www.cluster2010.org\">12th IEEE International Conference on Cluster Computing (Cluster) 2010<\/a>: <a href=\"http:\/\/www2.wmin.ac.uk\/getovv\/aacec10.html\" target=\"www2.wmin.ac.uk\/getovv\/aacec10.html\">1st Workshop on Application\/Architecture Co-design for Extreme-scale Computing (AACEC)<\/a><\/i>, pages 1-8, Hersonissos, Crete, Greece, September 20-24, 2010. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4244-8395-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CLUSTERWKSP.2010.5613113\" target=\"publication\">10.1109\/CLUSTERWKSP.2010.5613113<\/a>. <a href=\"javascript:showAbstract('This work focuses on tools for investigating algorithm performance at extreme scale with millions of concurrent threads and for evaluating the impact of future architecture choices to facilitate the co-design of high-performance computing (HPC) architectures and applications. The approach focuses on lightweight simulation of extreme-scale HPC systems with the needed amount of accuracy. The prototype presented in this paper is able to provide this capability using a parallel discrete event simulation (PDES), such that a Message Passing Interface (MPI) application can be executed at extreme scale, and its performance properties can be evaluated. The results of an initial prototype are encouraging as a simple hello world MPI program could be scaled up to 1,048,576 virtual MPI processes on a four-node cluster, and the performance properties of two MPI programs could be evaluated at up to 1,024 and 16,384 virtual MPI processes on the same system.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann10facilitating.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann10facilitating.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann10facilitating\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>George Ostrouchov, Thomas Naughton, Christian Engelmann, Geoffroy R. Vall&eacute;e, and Stephen L. Scott. <b>Nonparametric Multivariate Anomaly Analysis in Support of HPC Resilience<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.oerc.ox.ac.uk\/ieee\" target=\"www.oerc.ox.ac.uk\/ieee\">5th IEEE International Conference on e-Science (e-Science) 2009<\/a>: <a href=\"http:\/\/www.oerc.ox.ac.uk\/ieee\/workshops\/workshops\/computational-science\" target=\"www.oerc.ox.ac.uk\/ieee\/workshops\/workshops\/computational-science\">Workshop on Computational Science<\/a><\/i>, pages 80-85, Oxford, UK, December 9-11, 2009. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-1-4244-5946-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ESCIW.2009.5407992\" target=\"publication\">10.1109\/ESCIW.2009.5407992<\/a>. <a href=\"javascript:showAbstract('Large-scale computing systems provide great potential for scientific exploration. However, the complexity that accompanies these enormous machines raises challeges for both, users and operators. The effective use of such systems is often hampered by failures encountered when running applications on systems containing tens-of-thousands of nodes and hundreds-of-thousands of compute cores capable of yielding petaflops of performance. In systems of this size failure detection is complicated and root-cause diagnosis difficult. This paper describes our recent work in the identification of anomalies in monitoring data and system logs to provide further insights into machine status, runtime behavior, failure modes and failure root causes. It discusses the details of an initial prototype that gathers the data and uses statistical techniques for analysis.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ostrouchov09nonparametric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ostrouchov09nonparametric.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ostrouchov09nonparametric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Thomas Naughton, Wesley Bland, Geoffroy R. Vall&eacute;e, Christian Engelmann, and Stephen L. Scott. <b>Fault Injection Framework for System Resilience Evaluation &#8211; Fake Faults for Finding Future Failures<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.lrz-muenchen.de\/hpdc2009\" target=\"www.lrz-muenchen.de\/hpdc2009\">18th International Symposium on High Performance Distributed Computing (HPDC) 2009<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2009\" target=\"xcr.cenit.latech.edu\/resilience2009\">2nd Workshop on Resiliency in High Performance Computing (Resilience) 2009<\/a><\/i>, pages 23-28, Munich, Germany, June 9, 2009. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-60558-587-1. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1552526.1552530\" target=\"publication\">10.1145\/1552526.1552530<\/a>. <a href=\"javascript:showAbstract('As high-performance computing (HPC) systems increase in size and complexity they become more difficult to manage. The enormous component counts associated with these large systems lead to significant challenges in system reliability and availability. This in turn is driving research into the resilience of large scale systems, which seeks to curb the effects of increased failures at large scales by masking the inevitable faults in these systems. The basic premise being that failure must be accepted as a reality of large scale system and coped with accordingly through system resilience. A key component in the development and evaluation of system resilience techniques is having a means to conduct controlled experiments. A common method for performing such experiments is to generate synthetic faults and study the resulting effects. In this paper we discuss the motivation and our initial use of software fault injection to support the evaluation of resilience for HPC systems. We mention background and related work in the area and discuss the design of a tool to aid in fault injection experiments for both user-space (application-level) and system-level failures.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/naughton09fault.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/naughton09fault.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#naughton09fault\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Anand Tikotekar, Hong H. Ong, Sadaf Alam, Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, and Stephen L. Scott. <b>Performance Comparison of Two Virtual Machine Scenarios Using an HPC Application &#8211; A Case study Using Molecular Dynamics Simulations<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.csm.ornl.gov\/srt\/hpcvirt09\" target=\"www.csm.ornl.gov\/srt\/hpcvirt09\">3rd Workshop on System-level Virtualization for High Performance Computing (HPCVirt) 2009<\/a>, in conjunction with the <a href=\"http:\/\/www.eurosys.org\/2009\" target=\"www.eurosys.org\/2009\">4th ACM SIGOPS European Conference on Computer Systems (EuroSys) 2009<\/a><\/i>, pages 33-40, Nuremberg, Germany, March 30, 2009. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-60558-465-2. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1519138.1519143\" target=\"publication\">10.1145\/1519138.1519143<\/a>. <a href=\"javascript:showAbstract('Obtaining high flexibility to performance-loss ratio is a key challenge of today&amp;#39;s HPC virtual environment landscape. And while extensive research has been targeted at extracting more performance from virtual machines, the idea that whether novel virtual machine usage scenarios could lead to high flexibility Vs performance trade-off has received less attention. We, in this paper, take a step forward by studying and comparing the performance implications of running the Large-scale Atomic\/Molecular Massively Parallel Simulator (LAMMPS) application on two virtual machine configurations. First configuration consists of two virtual machines per node with 1 application process per virtual machine. The second configuration consists of 1 virtual machine per node with 2 processes per virtual machine. Xen has been used as an hypervisor and standard Linux as a guest virtual machine. Our results show that the difference in overall performance impact on LAMMPS between the two virtual machine configurations described above is around 3%. We also study the difference in performance impact in terms of each configuration's individual metrics such as CPU, I\/O, Memory, and interrupt\/context switches.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tikotekar09performance.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tikotekar09performance.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tikotekar09performance\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy R. Vall&eacute;e, Thomas Naughton, Hong H. Ong, Anand Tikotekar, Christian Engelmann, Wesley Bland, Ferrol Aderholt, and Stephen L. Scott. <b>Virtual System Environments<\/b>. In <i>Communications in Computer and Information Science: Proceedings of the <a href=\"http:\/\/www.dmtf.org\/svm08\" target=\"www.dmtf.org\/svm08\">2nd DMTF Academic Alliance Workshop on Systems and Virtualization Management: Standards and New Technologies (SVM) 2008<\/a><\/i>, pages 72-83, Munich, Germany, October 21-22, 2008. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-540-88707-2. ISSN 1865-0929. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-540-88708-9_7\" target=\"publication\">10.1007\/978-3-540-88708-9_7<\/a>. <a href=\"javascript:showAbstract('Distributed and parallel systems are typically managed with static settings: the operating system (OS) and the runtime environment (RTE) are specified at a given time and cannot be changed to fit an application`s needs. This means that every time application developers want to use their application on a new execution platform, the application has to be ported to this new environment, which may be expensive in terms of application modifications and developer time. However, the science resides in the applications and not in the OS or the RTE. Therefore, it should be beneficial to adapt the OS and the RTE to the application instead of adapting the applications to the OS and the RTE. This document presents the concept of Virtual System Environments (VSE), which enables application developers to specify and create a virtual environment that properly fits their application`s needs. For that four challenges have to be addressed: (i) definition of the VSE itself by the application developers, (ii) deployment of the VSE, (iii) system administration for the platform, and (iv) protection of the platform from the running VSE. We therefore present an integrated tool for the definition and deployment of VSEs on top of traditional and virtual (i.e., using system-level virtualization) execution platforms. This tool provides the capability to choose the degree of delegation for system administration tasks and the degree of protection from the application (e.g., using virtual machines). To summarize, the VSE concept enables the customization of the OS\/RTE used for the execution of application by users without compromising local system administration rules and execution platform protection constraints.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/vallee08virtual.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#vallee08virtual\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Anand Tikotekar, Geoffroy Vall&eacute;e, Thomas Naughton, Hong H. Ong, Christian Engelmann, and Stephen L. Scott. <b>An Analysis of HPC Benchmark Applications in Virtual Machine Environments<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/europar2008.caos.uab.es\" target=\"europar2008.caos.uab.es\">14th European Conference on Parallel and Distributed Computing (Euro-Par) 2008<\/a>: <a href=\"http:\/\/scilytics.com\/vhpc\" target=\"scilytics.com\/vhpc\">3rd Workshop on Virtualization in High-Performance Cluster and Grid Computing (VHPC) 2008<\/a><\/i>, pages 63-71, Las Palmas de Gran Canaria, Spain, August 26-29, 2008. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 978-3-642-00954-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-642-00955-6\" target=\"publication\">10.1007\/978-3-642-00955-6<\/a>. <a href=\"javascript:showAbstract('Virtualization technology has been gaining acceptance in the scientific community due to its overall flexibility in running HPC applications. It has been reported that a specific class of applications is better suited to a particular type of virtualization scheme or implementation. For example, Xen has been shown to perform with little overhead for compute-bound applications. Such a study, although useful, does not allow us to generalize conclusions beyond the performance analysis of that application which is explicitly executed. An explanation of why the generalization described above is difficult, may be due to the versatility in applications, which leads to different overheads in virtual environments. For example, two similar applications may spend disproportionate amount of time in their respective library code when run in virtual environments. In this paper, we aim to study such potential causes by investigating the behavior and identifying patterns of various overheads for HPC benchmark applications. Based on the investigation of the overhead profiles for different benchmarks, we aim to address questions such as: Are the overhead profiles for a particular type of benchmarks (such as compute-bound) similar or are there grounds to conclude otherwise?');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tikotekar08analysis.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tikotekar08analysis.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tikotekar08analysis\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Symmetric Active\/Active High Availability for High-Performance Computing System Services: Accomplishments and Limitations<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ens-lyon.fr\/LIP\/RESO\/ccgrid2008\" target=\"www.ens-lyon.fr\/LIP\/RESO\/ccgrid2008\">8th IEEE International Symposium on Cluster Computing and the Grid (CCGrid) 2008<\/a>: <a href=\"http:\/\/xcr.cenit.latech.edu\/resilience2008\" target=\"xcr.cenit.latech.edu\/resilience2008\">Workshop on Resiliency in High Performance Computing (Resilience) 2008<\/a><\/i>, pages 813-818, Lyon, France, May 19-22, 2008. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 978-0-7695-3156-4. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CCGRID.2008.78\" target=\"publication\">10.1109\/CCGRID.2008.78<\/a>. <a href=\"javascript:showAbstract('This paper summarizes our efforts over the last 3-4 years in providing symmetric active\/active high availability for high-performance computing (HPC) system services. This work paves the way for high-level reliability, availability and serviceability in extreme-scale HPC systems by focusing on the most critical components, head and service nodes, and by reinforcing them with appropriate high availability solutions. This paper presents our accomplishments in the form of concepts and respective prototypes, discusses existing limitations, outlines possible future work, and describes the relevance of this research to other, planned efforts.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann08symmetric2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann08symmetric2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann08symmetric2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xin Chen, Benjamin Eckart, Xubin (Ben) He, Christian Engelmann, and Stephen L. Scott. <b>An Online Controller Towards Self-Adaptive File System Availability and Performance<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2008\" target=\"xcr.cenit.latech.edu\/hapcw2008\">5th High Availability and Performance Workshop (HAPCW) 2008<\/a>, in conjunction with the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a><\/i>, Denver, CO, USA, April 3-4, 2008. <a href=\"javascript:showAbstract('At the present time, it can be a significant challenge to build a large-scale distributed file system that simultaneously maintains both high availability and high performance. Although many fault tolerance technologies have been proposed and used in both commercial and academic distributed file systems to achieve high availability, most of them typically sacrifice performance for higher system availability. Additionally, recent studies show that system availability and performance are related to the system workload. In this paper, we analyze the correlations among availability, performance, and workloads based on a replication strategy, and we discuss the trade off between availability and performance with different workloads. Our analysis leads to the design of an online controller that can dynamically achieve optimal performance and availability by tuning the system replication policy.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/chen08online.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/chen08online.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#chen08online\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Anand Tikotekar, Geoffroy Vall&eacute;e, Thomas Naughton, Hong H. Ong, Christian Engelmann, Stephen L. Scott, and Anthony M. Filippi. <b>Effects of Virtualization on a Scientific Application &#8211; Running a Hyperspectral Radiative Transfer Code on Virtual Machines<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.csm.ornl.gov\/srt\/hpcvirt08\" target=\"www.csm.ornl.gov\/srt\/hpcvirt08\">2nd Workshop on System-level Virtualization for High Performance Computing (HPCVirt) 2008<\/a>, in conjunction with the <a href=\"http:\/\/www.eurosys.org\/2008\" target=\"www.eurosys.org\/2008\">3rd ACM SIGOPS European Conference on Computer Systems (EuroSys) 2008<\/a><\/i>, pages 16-23, Glasgow, UK, March 31, 2008. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 978-1-60558-120-0. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1435452.1435455\" target=\"publication\">10.1145\/1435452.1435455<\/a>. <a href=\"javascript:showAbstract('The topic of system-level virtualization has recently begun to receive interest for high performance computing (HPC). This is in part due to the isolation and encapsulation offered by the virtual machine. These traits enable applications to customize their environments and maintain consistent software configurations in their virtual domains. Additionally, there are mechanisms that can be used for fault tolerance like live virtual machine migration. Given these attractive benefits to virtualization, a fundamental question arises, how does this effect my scientific application? We use this as the premise for our paper and observe a real-world scientific code running on a Xen virtual machine. We studied the effects of running a radiative transfer simulation, Hydrolight, on a virtual machine. We discuss our methodology and report observations regarding the usage of virtualization with this application.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tikotekar08effects.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tikotekar08effects.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tikotekar08effects\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Hong H. Ong, and Stephen L. Scott. <b>Middleware in Modern High Performance Computing System Architectures<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.iccs-meeting.org\/iccs2007\" target=\"www.iccs-meeting.org\/iccs2007\">7th International Conference on Computational Science (ICCS) 2007<\/a>, Part II: <a href=\"http:\/\/www.gup.uni-linz.ac.at\/cce2007\" target=\"www.gup.uni-linz.ac.at\/cce2007\">4th Special Session on Collaborative and Cooperative Environments (CCE) 2007<\/a><\/i>, pages 784-791, Beijing, China, May 27-30, 2007. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 3-5407-2585-5. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-540-72586-2_111\" target=\"publication\">10.1007\/978-3-540-72586-2_111<\/a>. <a href=\"javascript:showAbstract('A recent trend in modern high performance computing (HPC) system architectures employs lean compute nodes running a lightweight operating system (OS). Certain parts of the OS a well as other system software services are moved to service nodes in order to increase performance and scalability. This paper examines the impact of this HPC system architecture trend on HPC middleware software solutions, which traditionally equip HPC systems with advanced features, such as parallel and distributed programming models, appropriate system resource management mechanisms, remote application steering and user interaction techniques. Since the approach of keeping the compute node software stack small and simple is orthogonal to the middleware concept of adding missing OS features between OS and application, the role and architecture of middleware in modern HPC systems needs to be revisited. The result is a paradigm shift in HPC middleware design, where single middleware services are moved to service nodes, while runtime environments (RTEs) continue to reside on compute nodes.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07middleware.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07middleware.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07middleware\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Transparent Symmetric Active\/Active Replication for Service-Level High Availability<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ccgrid07.lncc.br\" target=\"ccgrid07.lncc.br\">7th IEEE International Symposium on Cluster Computing and the Grid (CCGrid) 2007<\/a>: <a href=\"http:\/\/www.lri.fr\/ fedak\/gp2pc-07\" target=\"www.lri.fr\/ fedak\/gp2pc-07\">7th International Workshop on Global and Peer-to-Peer Computing (GP2PC) 2007<\/a><\/i>, pages 755-760, Rio de Janeiro, Brazil, May 14-17, 2007. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2833-3. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/CCGRID.2007.116\" target=\"publication\">10.1109\/CCGRID.2007.116<\/a>. <a href=\"javascript:showAbstract('As service-oriented architectures become more important in parallel and distributed computing systems, individual service instance reliability as well as appropriate service redundancy becomes an essential necessity in order to increase overall system availability. This paper focuses on providing redundancy strategies using service-level replication techniques. Based on previous research using symmetric active\/active replication, this paper proposes a transparent symmetric active\/active replication approach that allows for more reuse of code between individual service-level replication implementations by using a virtual communication layer. Service- and client-side interceptors are utilized in order to provide total transparency. Clients and servers are unaware of the replication infrastructure as it provides all necessary mechanisms internally.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07transparent.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07transparent.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07transparent\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Hong H. Ong, Geoffroy R. Vall&eacute;e, and Thomas Naughton. <b>Configurable Virtualized System Environments for High Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.csm.ornl.gov\/srt\/hpcvirt07\" target=\"www.csm.ornl.gov\/srt\/hpcvirt07\">1st Workshop on System-level Virtualization for High Performance Computing (HPCVirt) 2007<\/a>, in conjunction with the <a href=\"http:\/\/www.eurosys.org\/2008\" target=\"www.eurosys.org\/2008\">2nd ACM SIGOPS European Conference on Computer Systems (EuroSys) 2007<\/a><\/i>, Lisbon, Portugal, March 20, 2007. <a href=\"javascript:showAbstract('Existing challenges for current terascale high performance computing (HPC) systems are increasingly hampering the development and deployment efforts of system software and scientific applications for next-generation petascale systems. The expected rapid system upgrade interval toward petascale scientific computing demands an incremental strategy for the development and deployment of legacy and new large-scale scientific applications that avoids excessive porting. Furthermore, system software developers as well as scientific application developers require access to large-scale testbed environments in order to test individual solutions at scale. This paper proposes to address these issues at the system software level through the development of a virtualized system environment (VSE) for scientific computing. The proposed VSE approach enables plug-and-play supercomputing through desktop-to-cluster-to-petaflop computer system-level virtualization based on recent advances in hypervisor virtualization technologies. This paper describes the VSE system architecture in detail, discusses needed tools for VSE system management and configuration, and presents respective VSE use case scenarios.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann07configurable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann07configurable.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann07configurable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Towards High Availability for High-Performance Computing System Services: Accomplishments and Limitations<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2006\" target=\"xcr.cenit.latech.edu\/hapcw2006\">4th High Availability and Performance Workshop (HAPCW) 2006<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.krellinst.org\" target=\"lacsi.krellinst.org\">7th Los Alamos Computer Science Institute (LACSI) Symposium 2006<\/a><\/i>, Santa Fe, NM, USA, October 17, 2006. <a href=\"javascript:showAbstract('During the last several years, our teams at Oak Ridge National Laboratory, Louisiana Tech University, and Tennessee Technological University focused on efficient redundancy strategies for head and service nodes of high-performance computing (HPC) systems in order to pave the way for high availability (HA) in HPC. These nodes typically run critical HPC system services, like job and resource management, and represent single points of failure and control for an entire HPC system. The overarching goal of our research is to provide high-level reliability, availability, and serviceability (RAS) for HPC systems by combining HA and HPC technology. This paper summarizes our accomplishments, such as developed concepts and implemented proof-of-concept prototypes, and describes existing limitations, such as performance issues, which need to be dealt with for production-type deployment.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06towards.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann06towards.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Li Ou, Xin Chen, Xubin (Ben) He, Christian Engelmann, and Stephen L. Scott. <b>Achieving Computational I\/O Effciency in a High Performance Cluster Using Multicore Processors<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2006\" target=\"xcr.cenit.latech.edu\/hapcw2006\">4th High Availability and Performance Workshop (HAPCW) 2006<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.krellinst.org\" target=\"lacsi.krellinst.org\">7th Los Alamos Computer Science Institute (LACSI) Symposium 2006<\/a><\/i>, Santa Fe, NM, USA, October 17, 2006. <a href=\"javascript:showAbstract('Cluster computing has become one of the most popular platforms for high-performance computing today. The recent popularity of multicore processors provides a flexible way to increase the computational capability of clusters. Although the system performance may improve with multicore processors in a cluster, I\/O requests initiated by multiple cores may saturate the I\/O bus, and furthermore increase the latency by issuing  multiple non-contiguous disk accesses. In this paper, we propose an asymmetric collective I\/O for multicore processors to improve multiple non-contiguous accesses. In our configuration, one core in each multicore processor is designated as the coordinator, and others serve as computing cores. The coordinator is responsible for aggregating I\/O operations from computing cores and submitting a contiguous request. The coordinator allocates contiguous memory buffers on behalf of other cores to avoid redundant data copies.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ou06achieving.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ou06achieving.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ou06achieving\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>RMIX: A Dynamic, Heterogeneous, Reconfigurable Communication Framework<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.iccs-meeting.org\/iccs2006\" target=\"www.iccs-meeting.org\/iccs2006\">6th International Conference on Computational Science (ICCS) 2006<\/a>, Part II: <a href=\"http:\/\/www.gup.uni-linz.ac.at\/cce2006\" target=\"www.gup.uni-linz.ac.at\/cce2006\">3rd Special Session on Collaborative and Cooperative Environments (CCE) 2006<\/a><\/i>, pages 573-580, Reading, UK, May 28-31, 2006. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 3-540-34381-4. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/11758525_77\" target=\"publication\">10.1007\/11758525_77<\/a>. <a href=\"javascript:showAbstract('RMIX is a dynamic, heterogeneous, reconfigurable communication framework that allows software components to communicate using various RMI\/RPC protocols, such as ONC RPC, Java RMI and SOAP, by facilitating dynamically loadable provider plug-ins to supply different protocol stacks. With this paper, we present a native (C-based), flexible, adaptable, multi-protocol RMI\/RPC communication framework that complements the Java-based RMIX variant previously developed by our partner team at Emory University. Our approach offers the same multi-protocol RMI\/RPC services and advanced invocation semantics via a C-based interface that does not require an object-oriented programming language. This paper provides a detailed description of our RMIX framework architecture and some of its features. It describes the general use case of the RMIX framework and its integration into the Harness metacomputing environment in the form of a plug-in.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06rmix.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann06rmix.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06rmix\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, Chokchai (Box) Leangsuksun, and Xubin (Ben) He. <b>Active\/Active Replication for Highly Available HPC System Services<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ares-conference.eu\/ares2006\" target=\"www.ares-conference.eu\/ares2006\">1st International Conference on Availability, Reliability and Security (ARES) 2006<\/a>: 1st International Workshop on Frontiers in Availability, Reliability and Security (FARES) 2006<\/i>, pages 639-645, Vienna, Austria, April 20-22, 2006. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2567-9. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ARES.2006.23\" target=\"publication\">10.1109\/ARES.2006.23<\/a>. <a href=\"javascript:showAbstract('Today`s high performance computing systems have several reliability deficiencies resulting in availability and serviceability issues. Head and service nodes represent a single point of failure and control for an entire system as they render it inaccessible and unmanageable in case of a failure until repair, causing a significant downtime. This paper introduces two distinct replication methods (internal and external) for providing symmetric active\/active high availability for multiple head and service nodes running in virtual synchrony. It presents a comparison of both methods in terms of expected correctness, ease-of-use and performance based on early results from ongoing work in providing symmetric active\/active high availability for two HPC system services (TORQUE and PVFS metadata server). It continues with a short description of a distributed mutual exclusion algorithm and a brief statement regarding the handling of Byzantine failures. This paper concludes with an overview of past and ongoing work, and a short summary of the presented research.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann06active.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann06active.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann06active\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Stephen L. Scott. <b>Concepts for High Availability in Scientific High-End Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2005\" target=\"xcr.cenit.latech.edu\/hapcw2005\">3rd High Availability and Performance Workshop (HAPCW) 2005<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.rice.edu\/symposium\/agenda_2005\" target=\"lacsi.rice.edu\/symposium\/agenda_2005\">6th Los Alamos Computer Science Institute (LACSI) Symposium 2005<\/a><\/i>, Santa Fe, NM, USA, October 11, 2005. <a href=\"javascript:showAbstract('Scientific high-end computing (HEC) has become an important tool for scientists world-wide to understand problems, such as in nuclear fusion, human genomics and nanotechnology. Every year, new HEC systems emerge on the market with better performance and higher scale. With only very few exceptions, the overall availability of recently installed systems has been lower in comparison to the same deployment phase of their predecessors. In contrast to the experienced loss of availability, the demand for continuous availability has risen dramatically due to the recent trend towards capability computing. In this paper, we analyze the existing deficiencies of current HEC systems and present several high availability concepts to counter the experienced loss of availability and to alleviate the expected impact on next-generation systems. We explain the application of these concepts to current and future HEC systems and list past and ongoing related research. This paper closes with a short summary of the presented work and a brief discussion of future efforts.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05concepts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05concepts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05concepts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Stephen L. Scott. <b>High Availability for Ultra-Scale High-End Scientific Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/coset.irisa.fr\" target=\"coset.irisa.fr\">2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2) 2005<\/a>, in conjunction with the <a href=\"http:\/\/ics05.csail.mit.edu\" target=\"ics05.csail.mit.edu\">19th ACM International Conference on Supercomputing (ICS) 2005<\/a><\/i>, Cambridge, MA, USA, June 19, 2005. <a href=\"javascript:showAbstract('Ultra-scale architectures for scientific high-end computing with tens to hundreds of thousands of processors, such as the IBM Blue Gene\/L and the Cray X1, suffer from availability deficiencies, which impact the efficiency of running computational jobs by forcing frequent checkpointing of applications. Most systems are unable to handle runtime system configuration changes caused by failures and require a complete restart of essential system services, such as the job scheduler or MPI, or even of the entire machine. In this paper, we present a flexible, pluggable and component-based high availability framework that expands today`s effort in high availability computing of keeping a single server alive to include all machines cooperating in a high-end scientific computing environment, while allowing adaptation to system properties and application needs.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Chokchai (Box) Leangsuksun, Venkata K. Munganuru, Tong Liu, Stephen L. Scott, and Christian Engelmann. <b>Asymmetric Active-Active High Availability for High-end Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/coset.irisa.fr\" target=\"coset.irisa.fr\">2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2) 2005<\/a>, in conjunction with the <a href=\"http:\/\/ics05.csail.mit.edu\" target=\"ics05.csail.mit.edu\">19th ACM International Conference on Supercomputing (ICS) 2005<\/a><\/i>, Cambridge, MA, USA, June 19, 2005. <a href=\"javascript:showAbstract('Linux clusters have become very popular for scientific computing at research institutions world-wide, because they can be easily deployed at a fairly low cost. However, the most pressing issues of today`s cluster solutions are availability and serviceability. The conventional Beowulf cluster architecture has a single head node connected to a group of compute nodes. This head node is a typical single point of failure and control, which severely limits availability and serviceability by effectively cutting off healthy compute nodes from the outside world upon overload or failure. In this paper, we describe a paradigm that addresses this issue using asymmetric active-active high availability. Our framework comprises of n + 1 head nodes, where n head nodes are active in the sense that they provide services to simultaneously incoming user requests. One standby server monitors all active servers and performs a fail-over in case of a detected outage. We present a prototype implementation based on a 2 + 1 solution and discuss initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/leangsuksun05asymmetric.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/leangsuksun05asymmetric.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#leangsuksun05asymmetric\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>A Lightweight Kernel for the Harness Metacomputing Framework<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.ipdps.org\/ipdps2005\" target=\"www.ipdps.org\/ipdps2005\">19th IEEE International Parallel and Distributed Processing Symposium (IPDPS) 2005<\/a>: <a href=\"http:\/\/www.cs.umass.edu\/ rsnbrg\/hcw2005\" target=\"www.cs.umass.edu\/ rsnbrg\/hcw2005\">14th Heterogeneous Computing Workshop (HCW) 2005<\/a><\/i>, Denver, CO, USA, April 4, 2005. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-2312-9. ISSN 1530-2075. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/IPDPS.2005.34\" target=\"publication\">10.1109\/IPDPS.2005.34<\/a>. <a href=\"javascript:showAbstract('Harness is a pluggable heterogeneous Distributed Virtual Machine (DVM) environment for parallel and distributed scientific computing. This paper describes recent improvements in the Harness kernel design. By using a lightweight approach and moving previously integrated system services into software modules, the software becomes more versatile and adaptable. This paper outlines these changes and explains the major Harness kernel components in more detail. A short overview is given of ongoing efforts in integrating RMIX, a dynamic heterogeneous reconfigurable communication framework, into the Harness environment as a new plug-in software module. We describe the overall impact of these changes and how they relate to other ongoing work.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann05lightweight.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann05lightweight.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann05lightweight\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, and George A. (Al) Geist. <b>High Availability through Distributed Control<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2004\" target=\"xcr.cenit.latech.edu\/hapcw2004\">2nd High Availability and Performance Workshop (HAPCW) 2004<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.rice.edu\/symposium\/agenda_2004\" target=\"lacsi.rice.edu\/symposium\/agenda_2004\">5th Los Alamos Computer Science Institute (LACSI) Symposium 2004<\/a><\/i>, Santa Fe, NM, USA, October 12, 2004. <a href=\"javascript:showAbstract('Cost-effective, flexible and efficient scientific simulations in cutting-edge research areas utilize huge high-end computing resources with thousands of processors. In the next five to ten years the number of processors in such computer systems will rise to tens of thousands, while scientific application running times are expected to increase further beyond the Mean-Time-To-Interrupt (MTTI) of hardware and system software components. This paper describes the ongoing research in heterogeneous adaptable reconfigurable networked systems (Harness) and its recent achievements in the area of high availability distributed virtual machine environments for parallel and distributed scientific computing. It shows how a distributed control algorithm is able to steer a distributed virtual machine process in virtual synchrony while maintaining consistent replication for high availability. It briefly illustrates ongoing work in heterogeneous reconfigurable communication frameworks and security mechanisms. The paper continues with a short overview of similar research in reliable group communication frameworks, fault-tolerant process groups and highly available distributed virtual processes. It closes with a brief discussion of possible future research directions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann04high.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann04high.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann04high\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Xubin (Ben) He, Li Ou, Stephen L. Scott, and Christian Engelmann. <b>A Highly Available Cluster Storage System using Scavenging<\/b>. In <i>Proceedings of the <a href=\"http:\/\/xcr.cenit.latech.edu\/hapcw2004\" target=\"xcr.cenit.latech.edu\/hapcw2004\">2nd High Availability and Performance Workshop (HAPCW) 2004<\/a>, in conjunction with the <a href=\"http:\/\/lacsi.rice.edu\/symposium\/agenda_2004\" target=\"lacsi.rice.edu\/symposium\/agenda_2004\">5th Los Alamos Computer Science Institute (LACSI) Symposium 2004<\/a><\/i>, Santa Fe, NM, USA, October 12, 2004. <a href=\"javascript:showAbstract('Highly available data storage for high-performance computing is becoming increasingly more critical as high-end computing systems scale up in size and storage systems are developed around network-centered architectures. A promising solution is to harness the collective storage potential of individual workstations much as we harness idle CPU cycles due to the excellent price\/performance ratio and low storage usage of most commodity workstations. For such a storage system, metadata consistency is a key issue assuring storage system availability as well as data reliability. In this paper, we present a decentralized metadata management scheme that improves storage availability without sacrificing performance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/he04highly.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/he04highly.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#he04highly\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and George A. (Al) Geist. <b>A Diskless Checkpointing Algorithm for Super-scale Architectures Applied to the Fast Fourier Transform<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.cs.msstate.edu\/ clade2003\" target=\"www.cs.msstate.edu\/ clade2003\">Challenges of Large Applications in Distributed Environments Workshop (CLADE) 2003<\/a>, in conjunction with the <a href=\"http:\/\/csag.ucsd.edu\/HPDC-12\" target=\"csag.ucsd.edu\/HPDC-12\">12th IEEE International Symposium on High Performance Distributed Computing (HPDC) 2003<\/a><\/i>, pages 47, Seattle, WA, USA, June 21, 2003. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 0-7695-1984-9. DOI <a href=\"http:\/\/dx.doi.org\/xpls\/abs_all.jsp?arnumber=4159902\" target=\"publication\">xpls\/abs_all.jsp?arnumber=4159902<\/a>. <a href=\"javascript:showAbstract('This paper discusses the issue of fault-tolerance in distributed computer systems with tens or hundreds of thousands of diskless processor units. Such systems, like the IBM Blue Gene\/L, are predicted to be deployed in the next five to ten years. Since a 100,000-processor system is going to be less reliable, scientific applications need to be able to recover from occurring failures more efficiently. In this paper, we adapt the present technique of diskless checkpointing to such huge distributed systems in order to equip existing scientific algorithms with super-scalable fault-tolerance. First, we discuss the method of diskless checkpointing, then we adapt this technique to super-scale architectures and finally we present results from an implementation of the Fast Fourier Transform that uses the adapted technique to achieve super-scale fault-tolerance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann03diskless.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann03diskless.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann03diskless\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Stephen L. Scott, and George A. (Al) Geist. <b>Distributed Peer-to-Peer Control in Harness<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.science.uva.nl\/events\/ICCS2002\" target=\"www.science.uva.nl\/events\/ICCS2002\">2nd International Conference on Computational Science (ICCS) 2002<\/a>, Part II: Workshop on Global and Collaborative Computing<\/i>, pages 720-727, Amsterdam, The Netherlands, April 21-24, 2002. <a href=\"http:\/\/www.springer.com\" target=\"www.springer.com\">Springer Verlag, Berlin, Germany<\/a>. ISBN 3-540-43593-X. ISSN 0302-9743. DOI <a href=\"http:\/\/dx.doi.org\/content\/l537ujfwt8yta2dp\" target=\"publication\">content\/l537ujfwt8yta2dp<\/a>. <a href=\"javascript:showAbstract('Harness is an adaptable fault-tolerant virtual machine environment for next-generation heterogeneous distributed computing developed as a follow on to PVM. It additionally enables the assembly of applications from plug-ins and provides fault-tolerance. This work describes the distributed control, which manages global state replication to ensure a high-availability of service. Group communication services achieve an agreement on an initial global state and a linear history of global state changes at all members of the distributed virtual machine. This global state is replicated to all members to easily recover from single, multiple and cascaded faults. A peer-to-peer ring network architecture and tunable multi-point failure conditions provide heterogeneity and scalability. Finally, the integration of the distributed control into the multi-threaded kernel architecture of Harness offers a fault-tolerant global state database service for plug-ins and applications.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann02distributed.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann02distributed.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann02distributed\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Christian Engelmann, Andrew Ayres, Stephen DeWitt, Michael J. Brim, and Brett Eiffert. Building Resilient Self-Driving Laboratories with the INTERSECT Federated Ecosystem. In Proceedings of the 39th International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2026: 8th Annual Workshop on Extreme-Scale Experiment-in-the-Loop Computing (XLOOP) 2026, Chicago, IL, USA, November 15, 2026. IEEE&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":16,"menu_order":2,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-96","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/96","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=96"}],"version-history":[{"count":11,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/96\/revisions"}],"predecessor-version":[{"id":1449,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/96\/revisions\/1449"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=96"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}