{"id":43,"date":"2023-08-21T08:00:31","date_gmt":"2023-08-21T08:00:31","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=43"},"modified":"2026-08-22T03:52:04","modified_gmt":"2026-08-22T03:52:04","slug":"peer-reviewed-conference-posters","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=43","title":{"rendered":"Peer-Reviewed Conference Posters"},"content":{"rendered":"<ol>\n<li>Swen Boehm, Craig A. Bridges, Patrick Widener, Terry Jones, Sheikh Ghafoor, Christian Engelmann, and Olga Kuchar. <b>From Automated Experiments and Simulations to Reusable Scientific Evidence<\/b>. Poster at the <a href=\"http:\/\/www.montereydataconference.org\" target=\"www.montereydataconference.org\">Monterey Data Conference<\/a>, Monterey, CA, USA, August 24-26, 2026. <a href=\"javascript:showAbstract('Scientific research increasingly depends on automated platforms and heterogeneous instruments that generate rich, multi-modal data. Yet fragmentation across domain-specific systems, proprietary formats, and siloed repositories impedes data discovery, integration, and reuse. Without systematic semantic representation and comprehensive provenance, high-quality experimental data loses much of its scientific value. The INTERSECT Scientific data layer (SDL) provides an integrated, ontology-driven ecosystem that connects scientific platforms, workflows, and data management services into a coherent whole. Built on a system-of-systems architecture and grounded in Linked Data Platform (LDP) principles, the SDL enables modular integration of diverse services while preserving interoperability across scientific domains.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm26automated.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#boehm26automated\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Swen Boehm, Terry Jones, Patrick Widener, Christian Engelman, and Olga Kuchar. <b>INTERSECT Scientific Data Layer: A Federated, Modular Framework for Scientific Data Management<\/b>. Poster at the <a href=\"http:\/\/data-science.llnl.gov\/d3\" target=\"data-science.llnl.gov\/d3\">Department of Energy (DOE) Data Days (D3) Workshop<\/a>, Chantilly, VA, USA, March 3-5, 2026. <a href=\"javascript:showAbstract('Scientific research increasingly depends on automated platforms and heterogeneous instruments that generate rich, multi-modal data. Yet fragmentation across domain-specific systems, proprietary formats, and siloed repositories impedes data discovery, integration, and reuse. Without systematic semantic representation and comprehensive provenance, high-quality experimental data loses much of its scientific value.  The INTERSECT Scientific data layer (SDL) provides an integrated, ontology-driven ecosystem that connects scientific platforms, workflows, and data management services into a coherent whole. Built on a system-of-systems architecture and grounded in Linked Data Platform (LDP) principles, the SDL enables modular integration of diverse services while preserving interoperability across scientific domains. Core ontologies such as SSN\/SOSA for sensor and observation modeling, DCAT for resource cataloging, and PROV-O for provenance tracking provide a semantic backbone that ensures all entities - data, instruments, workflows, and results - are described in a machine-actionable, reusable way.  The SDL offers semantic-first design, a microservices foundation, separation of concerns, and content negotiation:  - Semantic-First Design: RDF is the native data model, not   an auxiliary export format. Semantic richness is preserved   throughout the data lifecycle, from instrumental observations   through processing pipelines to publication, enabling FAIR   data by design. - Microservices Foundation: Modular, independently deployable   services (Catalog Service, Storage Service, Repository   Service, Registry Service) coordinate through shared   semantic libraries and standard ontologies, solving the   distributed consistency challenge inherent in semantic   systems. - Separation of Concerns: Semantic metadata (RDF triples in   triple stores) is decoupled from data artifacts (files in   object storage) with URIs providing semantic linking. This   enables independent scaling of metadata management and   storage infrastructure while maintaining coherent provenance   relationships. - Content Negotiation: Services accept and return data in   multiple RDF serializations (Turtle, JSON-LD, RDF\/XML) and   domain-specific formats (CSV, HDF5, instrument formats),   supporting diverse tools and workflows while maintaining    semantic consistency.  The SDL natively implements FAIR principles through semantic-first architecture. Persistent URIs and SPARQL endpoints enable discovery via machine-readable metadata (findable). Standard HTTP protocols and LDP containers support predictable REST-like access patterns (accessible). Composed W3C ontologies ensure semantic compatibility across domains (interoperable). Comprehensive end-to-end provenance and structured metadata make datasets suitable for both human researchers and AI systems (reusable).');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/boehm26intersect.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#boehm26intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Swen Boehm, Michael Brim, Jack Lange, Thomas Naughton, Patrick Widener, Ben Mintz, and Rohit Srivastava. <b>INTERSECT: The Open Federated Architecture for the Laboratory of the Future<\/b>. Poster at the <a href=\"http:\/\/icpp23.sci.utah.edu\/\" target=\"icpp23.sci.utah.edu\/\">52nd International Conference on Parallel Processing (ICPP) 2023<\/a>, Salt Lake City, UT, USA, August 7-10, 2023. <a href=\"javascript:showAbstract('The open Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) architecture connects scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing center to enable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery and evaluation. Its a novel approach consists of science use case design patterns, a system of systems architecture, and a microservice architecture.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann23intersect.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann23intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Mohit Kumar. <b>Resilience Design Patterns: A Structured Modeling Approach of Resilience in Computing Systems<\/b>. Poster at the <a href=\"http:\/\/www.bnl.gov\/modsim2022\" target=\"www.bnl.gov\/modsim2022\">Workshop on Modeling and Simulation of Systems and Applications (ModSim) 2022<\/a>, Seattle, WA, USA, August 10-12, 2022. <a href=\"javascript:showAbstract('Resilience to faults, errors, and failures in extreme-scale high-performance computing (HPC) systems is a critical challenge. Resilience design patterns (Figure 1) offer a new, structured hardware\/software design approach for improving resilience by identifying and evaluating repeatedly occurring resilience problems and coordinating corresponding solutions. Initial work identified and formalized these patterns and developed a proof-of-concept prototype to demonstrate portable resilience. This recent work created performance, reliability, and availability models for each of the identified 15 structural resilience design patterns and a modeling tool that allows (1) exploring the performance, reliability, and availability of each pattern, and (2) investigating the trade-offs be-tween patterns and pattern combinations.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann22resilience.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann22resilience\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Rizwan Ashraf, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>Real-Time Assessment of Supercomputer Status by a Comprehensive Informative Metric through Streaming Processing<\/b>. Poster at the  <a href=\"http:\/\/cci.drexel.edu\/bigdata\/bigdata2018\" target=\"cci.drexel.edu\/bigdata\/bigdata2018\">6th IEEE International Conference on Big Data (BigData) 2018<\/a>,  Seattle, WA, USA, December 10-13, 2018. <a href=\"javascript:showAbstract('Supercomputers are complex systems used to simulate, understand and solve real-world problems. In order to operate these systems efficiently and for the purpose of their maintainability, an accurate, concise, and timely determination of system status is crucial for its users and operators. However, this determination is challenging due to intricately connected heterogeneous software and hardware components, and due to sheer scale of such machines. In this poster, we demonstrate work-in-progress towards realization of a real-time monitoring framework for the 18,688-node Titan supercomputer at Oak Ridge Leadership Computing Facility (OLCF). Toward this end, we discuss the use of metrics which present a one-dimensional view of the system generating various types of information from 1000s of components and utilization statistics from 100s of user applications in near real-time. We demonstrate the efficacy of these metrics to understand and visualize raw log data generated by the system which otherwise may compose of 1000s of dimensions. We also demonstrate the architecture of proposed real-time stream processing framework which integrates, processes, analyzes, visualizes and stores system log data from an array of system components..');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18realtime.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hui18realtime\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Yawei Hui, Byung Hoon (Hoony) Park, and Christian Engelmann. <b>A Comprehensive Informative Metric for Summarizing HPC System Status<\/b>. Poster at the <a href=\"http:\/\/ldav.org\" target=\"ldav.org\">8th IEEE Symposium on Large Data Analysis and   Visualization<\/a> in conjunction with the   <a href=\"http:\/\/ieeevis.org\/year\/2018\" target=\"ieeevis.org\/year\/2018\">8th IEEE Vis 2018<\/a>,  Berlin, Germany, October 21, 2018. <a href=\"javascript:showAbstract('It remains a major challenge to effectively summarize and visualize in a comprehensive form the status of a complex computer system, such as the Titan supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). In the ongoing research highlighted in this poster, we present system information entropy (SIE), a newly developed system metric that leverages the powers of traditional machine learning techniques and information theory. By compressing the multi-variant multi-dimensional event information recorded during the operation of the targeted system into a single time series of SIE, we demonstrate that the historical system status can be sensitively summarized in form of SIE and visualized concisely and comprehensively.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/hui18comprehensive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#hui18comprehensive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Rizwan Ashraf. <b>Modeling and Simulation of Extreme-Scale Systems for Resilience by Design<\/b>. Poster at the <a href=\"http:\/\/www.bnl.gov\/modsim2018\" target=\"www.bnl.gov\/modsim2018\">Workshop on Modeling and Simulation of Systems and Applications<\/a>, Seattle, WA, USA, August 15-17, 2018. <a href=\"javascript:showAbstract('Resilience is a serious concern for extreme-scale high-performance computing (HPC). While the HPC community has developed various resilience solutions, the solution space remains fragmented. We created a structured approach to the design, evaluation and optimization of HPC resilience using the concept of design patterns. A design pattern describes a generalized solution to a repeatedly occurring problem. We identified the commonly occurring problems and solutions used to deal with faults, errors and failures in HPC systems. Each well-known solution that addresses a specific resilience challenge is described in the form of a design pattern. We developed a resilience design pattern specification, language and catalog, which can be used by system architects, system software and library developers, application programmers, as well as users and operators as essential building blocks when designing and deploying resilience solutions. The resilience design pattern approach provides a unique opportunity for design space exploration. As each resilience solution is abstracted as a pattern and each solution&amp;#39;s properties are defined by pattern parameters, vertical and horizontal pattern compositions can describe the resilience capabilities of an entire HPC system. This permits the investigation of beneficial or counterproductive interactions between patterns and of the performance, resilience, and power consumption trade-off between different pattern parameters and compositions. The ultimate goal is to make resilience an integral part of the HPC hardware\/software ecosystem by coordinating the various existing resilience solutions in a design space exploration process, such that the burden for providing resilience is on the system by design and not on the user as an afterthought. We are in the early stages of developing a novel design space exploration tool that enables this investigation using modeling and simulation. We developed performance and resilience models for each resilience design pattern. We also leverage results from the Catalog project, a collaborative effort between Oak Ridge National Laboratory, Argonne National Laboratory and Lawrence Livermore National Laboratory that developed models of the faults, errors and failures in today's HPC systems. We also leverage recent results from the same project by Lawrence Livermore National Laboratory in application reliability patterns. The planned research extends and combines this work to model the performance, resilience, and power consumption of an entire HPC system, initially at node-level granularity, and to simulate the dynamic interactions between deployed resilience solutions and the rest of the system. In the next iteration, finer-grain modeling and simulation, such as at the computational unit level, is used to increase accuracy. This work leverages the experience of the investigators in parallel discrete event simulation of extreme-scale systems, such as the Extreme-scale Simulator (xSim). The current state of the art in resilience modeling and simulation is fragmented as well. There is currently no such design space exploration tool. Instead, each resilience solution is typically investigated separately. There is only a small amount of work on multi-resilience solutions, including by the investigators. While there is work in investigating the performance\/resilience trade-off space, there is almost no work in including power consumption.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18modeling2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann18modeling2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Onkar Patil, Saurabh Hukerikar, Frank Mueller, and Christian Engelmann. <b>Exploring Use Cases for Non-Volatile Memories in Support of HPC Resilience<\/b>. Poster at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">30th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2017<\/a>, Denver, CO, USA, November 12-17, 2017. <a href=\"javascript:showAbstract('Improving resilience and creating resilient architectures is one of the major goals of exascale computing. With the advent of Non-volatile memory technologies, memory architectures with persistent memory regions will be a significant part of future architectures. There is potential to use them in more than one way to benefit different applications. We look to take advantage of this technology to enable more fine-grained and novel methodology that will improve resilience and efficiency of exascale applications. We have developed three modes of memory usage for persistent memory to enable efficient checkpointing in HPC applications. We have developed a simple API that is evaluated with the DGEMM benchmark on a 16-node cluster with independent SSDs on every node. Our aim is to build on this work and enable static and dynamic runtime systems that will inherently make the HPC applications more fault-tolerant and resistant to errors.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/patil17exploring.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#patil17exploring\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Frank Mueller, Christian Engelmann, Rolf Riesen, and Kurt Ferreira. <b>Detection and Correction of Silent Data Corruption for Large-Scale High-Performance Computing<\/b>. Poster at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">24th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2011<\/a>, Seattle, WA, USA, November 12-18, 2011. <a href=\"javascript:showAbstract('Faults have become the norm rather than the exception for high-end computing on clusters with 10s\/100s of thousands of cores. Exacerbating this situation, some of these faults will not be detected, manifesting themselves as silent errors that will corrupt memory while applications continue to operate and report incorrect results. This poster introduces RedMPI, an MPI library which resides in the MPI profiling layer. RedMPI is capable of both online detection and correction of soft errors that occur in MPI applications without requiring any modifications to the application source. By providing redundancy, RedMPI is capable of transparently detecting corrupt messages from MPI processes that become faulted during execution. Furthermore, with triple redundancy RedMPI additionally &amp;#34;votes&amp;#34; out MPI messages of a faulted process by replacing corrupted results with corrected results from unfaulted processes. We present an experimental evaluation of RedMPI on an assortment of applications to demonstrate the effectiveness of this approach.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#fiala11detection\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>David Fiala, Kurt Ferreira, Frank Mueller, and Christian Engelmann. <b>A Tunable, Software-based DRAM Error Detection and Correction Library for HPC<\/b>. Poster at the <a href=\"http:\/\/sc11.supercomputing.org\" target=\"sc11.supercomputing.org\">24th IEEE\/ACM International Conference on High Performance  Computing, Networking, Storage and Analysis (SC) 2011<\/a>, Seattle, WA, USA, November 12-18, 2011. <a href=\"javascript:showAbstract('Proposed exascale systems will present a number of considerable resiliency challenges. In particular, DRAM soft-errors, or bit-flips, are expected to greatly increase due to the increased memory density of these systems. Current hardware-based fault-tolerance methods will be unsuitable for addressing the expected soft error frequency rate. As a result, additional software will be needed to address this challenge. In this paper we introduce LIBSDC, a tunable, transparent silent data corruption detection and correction library for HPC applications. LIBSDC provides comprehensive SDC protection for program memory by implementing on-demand page integrity verification by utilizing the MMU. Experimental  benchmarks with Mantevo HPCCG show that once tuned, LIBSDC is able to achieve SDC protection with less than 100% overhead of resources.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#fiala11tunable2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, George Ostrouchov, Chokchai (Box) Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller, Chao Wang, Arun B. Nagarajan, and Jyothish Varma. <b>A Tunable Holistic Resiliency Approach for High-Performance Computing Systems<\/b>. Poster at the <a href=\"http:\/\/institute.lanl.gov\/resilience\/conferences\/2009\" target=\"institute.lanl.gov\/resilience\/conferences\/2009\">National HPC Workshop on Resilience 2009<\/a>, Arlington, VA, USA, August 12-14, 2009. <a href=\"javascript:showAbstract('In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation extreme-scale high-performance computing (HPC) systems. This poster describes our past and ongoing efforts in novel fault resilience technologies for HPC. Presented work includes proactive fault resilience techniques, system and application reliability models and analyses, failure prediction, transparent process- and virtual-machine-level migration, and trade-off models for combining preemptive migration with checkpoint\/restart. This poster summarizes our work and puts all individual technologies into context with a proposed holistic fault resilience framework.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott09tunable2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott09tunable2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, Christian Engelmann, and Hong H. Ong. <b>System-level Virtualization for for High-Performance Computing<\/b>. Poster at the <a href=\"http:\/\/institute.lanl.gov\/resilience\/conferences\/2009\" target=\"institute.lanl.gov\/resilience\/conferences\/2009\">National HPC Workshop on Resilience 2009<\/a>, Arlington, VA, USA, August 12-14, 2009. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel system software solutions for providing a virtual system environment (VSE) for next-generation extreme-scale high-performance computing (HPC) systems and beyond. The poster showcases results of developed proof-of-concept implementations and performed theoretical analyses, outlines planned research and development activities, and presents respective initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott09systemlevel.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott09systemlevel\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, George Ostrouchov, Chokchai (Box) Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller, Chao Wang, Arun B. Nagarajan, and Jyothish Varma. <b>A Tunable Holistic Resiliency Approach for High-Performance Computing Systems<\/b>. Poster at the <a href=\"http:\/\/ppopp09.rice.edu\" target=\"ppopp09.rice.edu\">14th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP) 2009<\/a>, Raleigh, NC, USA, February 14-18, 2009. <a href=\"javascript:showAbstract('In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation extreme-scale high-performance computing (HPC) systems. This poster describes our past and ongoing efforts in novel fault resilience technologies for HPC. Presented work includes proactive fault resilience techniques, system and application reliability models and analyses, failure prediction, transparent process- and virtual-machine-level migration, and trade-off models for combining preemptive migration with checkpoint\/restart. This poster summarizes our work and puts all individual technologies into context with a proposed holistic fault resilience framework.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott09tunable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott09tunable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>George A. (Al) Geist, Christian Engelmann, Jack J. Dongarra, George Bosilca, Magdalena M. S&#322;awi&#324;ska, and Jaros&#322;aw K. S&#322;awi&#324;ski. <b>The Harness Workbench: Unified and Adaptive Access to Diverse High-Performance Computing Platforms<\/b>. Poster at the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a>, Denver, CO, USA, March 30 &#8211; April 5, 2008. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel software solutions for providing unified and adaptive access to diverse high-performance computing (HPC) platforms. The poster showcases developed proof-of-concept implementations of tools and mechanisms that simplify scientific application development and deployment tasks, such that only minimal adaptation is needed when moving from one HPC system to another or after HPC system upgrades.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/geist08harness.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#geist08harness\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Christian Engelmann, Hong H. Ong, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, George Ostrouchov, Chokchai (Box) Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller, Chao Wang, Arun B. Nagarajan, Jyothish Varma, Xubin (Ben) He, Li Ou, and Xin Chen. <b>Resiliency for High-Performance Computing Systems<\/b>. Poster at the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a>, Denver, CO, USA, March 30 &#8211; April 5, 2008. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel system software solutions for providing high-level reliability, availability and serviceability (RAS) for next-generation extreme-scale high-performance computing (HPC) systems and beyond. The poster showcases results of developed proof-of-concept implementations and performed theoretical analyses, outlines planned research and development activities, and presents respective initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott08resiliency.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott08resiliency\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Stephen L. Scott, Geoffroy R. Vall&eacute;e, Thomas Naughton, Anand Tikotekar, Christian Engelmann, and Hong H. Ong. <b>System-level Virtualization for for High-Performance Computing<\/b>. Poster at the <a href=\"http:\/\/www.hpcsw.org\" target=\"www.hpcsw.org\">1st High-Performance Computer Science Week (HPCSW) 2008<\/a>, Denver, CO, USA, March 30 &#8211; April 5, 2008. <a href=\"javascript:showAbstract('This poster summarizes our past and ongoing research and development efforts in novel system software solutions for providing a virtual system environment (VSE) for next-generation extreme-scale high-performance computing (HPC) systems and beyond. The poster showcases results of developed proof-of-concept implementations and performed theoretical analyses, outlines planned research and development activities, and presents respective initial results.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/scott08systemlevel.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#scott08systemlevel\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Poster\" height=\"10pt\"> Poster, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Swen Boehm, Craig A. Bridges, Patrick Widener, Terry Jones, Sheikh Ghafoor, Christian Engelmann, and Olga Kuchar. From Automated Experiments and Simulations to Reusable Scientific Evidence. Poster at the Monterey Data Conference, Monterey, CA, USA, August 24-26, 2026. Swen Boehm, Terry Jones, Patrick Widener, Christian Engelman, and Olga Kuchar. INTERSECT Scientific Data Layer: A Federated, Modular&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":16,"menu_order":3,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-43","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/43","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=43"}],"version-history":[{"count":5,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/43\/revisions"}],"predecessor-version":[{"id":1444,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/43\/revisions\/1444"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=43"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}