{"id":2,"date":"2026-09-09T08:00:34","date_gmt":"2026-09-09T08:00:34","guid":{"rendered":"http:\/\/christian-engelmann.de\/?page_id=2"},"modified":"2026-09-09T21:57:37","modified_gmt":"2026-09-09T21:57:37","slug":"sample-page","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/","title":{"rendered":"About Me"},"content":{"rendered":"<p align=\"center\"><i><b>Extreme-Scale Computing | Fault Resilience | HW\/SW Co-Design Tools | Computing Continuum | Autonomous Experiments<\/b><\/i><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" src=\"images\/christian_engelmann.png\" hspace=\"10\" vspace=\"0\" height=\"145\" width=\"145\" align=\"left\" style=\"padding-right:5pt;\">Dr. Christian Engelmann is a <a href=\"https:\/\/www.ornl.gov\/staff-profile\/christian-engelmann\" target=\"www.ornl.gov_staff-profile_christian-engelmann\">Distinguished Computer Scientist<\/a> and the <a href=\"https:\/\/www.ornl.gov\/group\/isf\" target=\"www.ornl.gov_group_isf\">Intelligent Systems and Facilities Research Group<\/a> Leader at <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\">Oak Ridge National Laboratory (ORNL)<\/a>, the US Department of Energy&#8217;s (DOE) largest multiprogram science and technology laboratory with an annual budget of $2.6 billion and 7,000+ staff. He has more than 25 years experience in software research and development for extreme-scale high-performance computing (HPC) systems. Dr. Engelmann&#8217;s research solves computer science challenges in HPC software, such as scalability, dependability, and interoperability.<\/p>\n<p>Dr. Engelmann&#8217;s primary expertise is in <a href=\"?page_id=430#resilience\">HPC resilience<\/a>, i.e., efficiency and correctness in the presence of faults, errors, and failures. He is a leading HPC resilience expert and was a member of the DOE Technical Council on HPC Resilience 2013-15. He received the 2015 DOE Early Career Award for research in <a href=\"?page_id=475\">resilience design patterns<\/a>. Dr. Engelmann&#8217;s secondary expertise is in <a href=\"?page_id=570\">system software for the instrument-to-edge-to-Cloud-to-center computing continuum<\/a>, enabling science breakthroughs with autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence (AI) driven design, discovery and evaluation. He further has expertise in <a href=\"?page_id=433\">lightweight simulation of future-generation extreme-scale supercomputers<\/a>, studying the impact of hardware\/software properties on performance and resilience for application-architecture co-design. Dr. Engelmann is also an expert in operating system and runtime software for parallel and distributed systems.<\/p>\n<p>Dr. Engelmann earned a Dipl.-Ing. (FH) in Computer Systems Engineering from the University of Applied Sciences Berlin, Germany, and a M.Sc. in Computer Science from the University of Reading, UK, both in 2001 as conjoint degrees, and a Ph.D. in Computer Science from the University of Reading in 2008. He is a Senior Member of the Association for Computing Machinery (ACM) and the Institute of Electrical and Electronics Engineers (IEEE). In 2025, was recognized as a Distinguished Contributor of the IEEE Computer Society. He is also a Member of the Society for Industrial and Applied Mathematics (SIAM) and the Advanced Computing Systems Association (USENIX). <\/p>\n<p align=\"center\"><a href=\"http:\/\/www.linkedin.com\/in\/christianengelmann\" target=\"www.linkedin.com_in_christianengelmann\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"images\/linkedin.gif\" width=\"105\" height=\"15\" border=\"0\" hspace=\"0\" vspace=\"0\" alt=\"View Christian Engelmann's profile on LinkedIn\" style=\"padding-top:4pt; vertical-align: middle;\"><\/a> | <a href=\"https:\/\/www.xing.com\/profile\/Christian_Engelmann7\" target=\"www.xing.com_profile_Christian_Engelmann7\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"images\/xing.png\" width=\"40\" height=\"25\" border=\"0\" hspace=\"0\" vspace=\"0\" style=\"vertical-align: middle;\"><\/a> | <a href=\"https:\/\/scholar.google.com\/citations?user=99-rbwsAAAAJ\" target=\"scholar.google.com\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"images\/google-scholar.jpg\" width=\"69\" height=\"30\" border=\"0\" hspace=\"0\" vspace=\"0\" alt=\"View Christian Engelmann's profile on Google Scholar\" style=\"vertical-align: middle;\"><\/a> | <a href=\"https:\/\/dblp.org\/pid\/71\/4514\" target=\"dblp.org_pid_71_4514\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"images\/dblp.png\" width=\"60\" height=\"30\" border=\"0\" hspace=\"0\" vspace=\"0\" alt=\"DBLP: Christian Engelmann\" style=\"vertical-align: middle;\"><\/a> | <span style=\"white-space: nowrap;\">Scopus ID: <a href=\"https:\/\/www.scopus.com\/authid\/detail.uri?authorId=18037364000\" target=\"www.scopus.com\" rel=\"noopener\">18037364000<\/a><\/span> | <span style=\"white-space: nowrap;\"><a href=\"https:\/\/orcid.org\/0000-0003-4365-6416\" target=\"orcid.widget\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/orcid.org\/sites\/default\/files\/images\/orcid_16x16.png\" width=\"16\" height=\"16\" border=\"0\" alt=\"ORCID iD icon\" style=\"vertical-align: middle; margin-right:.25em;\">orcid.org\/0000-0003-4365-6416<\/a><\/span> | <a href=\"https:\/\/github.com\/engelmannc\" target=\"github_engelmannc\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"images\/github.png\" width=\"76\" height=\"18\" border=\"0\" alt=\"GitHub icon\" style=\"vertical-align: middle;\"><\/a><\/p>\n<p align=\"center\"><i><span style=\"white-space: nowrap;\">Contact: <a href=\"mailto:engelmannc@computer.org\">engelmannc@computer.org<\/a><\/span> | <span style=\"white-space: nowrap;\">2-page biography: <a href=\"engelmann.pdf\"><img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\" hspace=\"0\"><\/a><\/span> | <span style=\"white-space: nowrap;\">Resume: Available upon <a href=\"mailto:engelmannc@computer.org\">request<\/a><\/span><\/i><\/p>\n<h4>Ongoing Projects<\/h4>\n<p><b>2025-&#8230;:<\/b> The <a href=\"https:\/\/www.energy.gov\/undersecretaryforscience\/genesis-mission\/modcon-transformational-ai-and-data\" target=\"www.energy.gov_undersecretaryforscience_genesis-mission_modcon-transformational-ai-and-data\">Transformational AI Model Consortium<\/a>: Creating the Data Broker Standards for the <a href=\"https:\/\/genesis.energy.gov\/\" target=\"genesis.energy.gov\">DOE Genesis Mission<\/a><br \/>\n<b>2025-&#8230;:<\/b> The <a href=\"https:\/\/amsc.energy.gov\/\" target=\"amsc.energy.gov\">American Science Cloud (AmSC)<\/a>: Designing the Data Service architecture and APIs for the <a href=\"https:\/\/genesis.energy.gov\/\" target=\"genesis.energy.gov\">DOE Genesis Mission<\/a><br \/>\n<b>2024-&#8230;:<\/b> The <a href=\"?page_id=1120\">Resilient Federated Ecosystem for Self-Driving Laboratories<\/a> project creates an error- and failure-resilient federated ecosystem for instrument science, enabling reliable autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence driven design, discovery, and evaluation.<br \/>\n<b>2024-&#8230;:<\/b> The <a href=\"?page_id=1188\">Privacy-Preserving Federated Learning for Science: Building Sustainable and Trustworthy Foundation Models<\/a> project creates develops efficient communication, memory, and energy optimization techniques for federated learning algorithms, particularly for large-scale foundation models, while ensuring fairness and incentivizing participation.\n<\/p>\n<h4>Recently In the News<\/h4>\n<p><b>2025-12:<\/b> IEEE Computer Socienty. The 2025 Class of the <a href=\"https:\/\/www.computer.org\/membership\/distinguished-contributors\" target=\"energy-department-advances-investments-ai-science\">Distinguished Contributor Recognition Program<\/a> recognizes members for technical contributions to the computing profession, computing community, and humanity.<br \/>\n<b>2025-12-10:<\/b> DOE. <a href=\"https:\/\/www.energy.gov\/articles\/energy-department-advances-investments-ai-science\" target=\"energy-department-advances-investments-ai-science\">Energy Department Advances Investments in AI for Science<\/a>.<br \/>\n<b>2025-07-08:<\/b> DOE Advanced Scientific Computing Research. <a href=\"https:\/\/science.osti.gov\/-\/media\/ascr\/pdf\/facilities\/ALCC\/ALCC_Factsheets_2025.pdf\" target=\"ALCC_Factsheets_2025\">1.1 million supercomputer node-hours awarded to Privacy-Preserving Federated Learning for Foundation Models<\/a>.<br \/>\n<b>2025-04-16:<\/b> ORNL Review. <a href=\"https:\/\/www.ornl.gov\/labsofthefuture\" target=\"www.ornl.gov_labsofthefuture\">Creating the lab of the future: INTERSECT unites AI and automation to revolutionize scientific discovery<\/a>.<br \/>\n<b>2024-10-15:<\/b> ORNL News. <a href=\"https:\/\/www.ornl.gov\/news\/ornl-projects-included-67-million-doe-ai-science-research\" target=\"www.ornl.gov_news_ornl-projects-included-67-million-doe-ai-science-research\">New ORNL projects included in $67 million from DOE for AI in science research<\/a>.\n<\/p>\n<h4>Latest Peer-Reviewed Publications<\/h4>\n<ol>\n<li>C. Engelmann, A. Ayres, S. DeWitt, M. J. Brim, and B. Eiffert. <b>Building Resilient Self-Driving Laboratories with the INTERSECT Federated Ecosystem<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc26.supercomputing.org\" target=\"sc26.supercomputing.org\">39th International Conference on High Performance Computing, Networking, Storage and Analysis (SC) Workshops 2026<\/a>: <a href=\"http:\/\/wordpress.cels.anl.gov\/xloop-2026\/\" target=\"wordpress.cels.anl.gov\/xloop-2026\/\">8th Annual Workshop on Extreme-Scale Experiment-in-the-Loop Computing (XLOOP) 2026<\/a><\/i>, November, 2026. To appear. <a href=\"javascript:showAbstract('Failure resilience in federated ecosystems for instrument science presents a critical challenge. Failures disrupt experiments and make them potentially useless, wasting valuable resources and creating setbacks. Oak Ridge National Laboratory&amp;#39;s Self-driven Experiments for Science \/ Interconnected Science Ecosystem (INTERSECT) offers a federated ecosystem for instrument science, enabling autonomous experiments, self-driving laboratories, smart manufacturing, and AI-driven design, discovery, and evaluation. This paper documents the recent advances in creating a resilient INTERSECT ecosystem. The proposed solution includes a resilient architecture with resilience design patterns, a resilient system of systems (SoS) architecture, and a resilient microservices architecture; and a resilient software development kit with reliable service communication and asynchronous and synchronous failure detection and notification. The resilience capabilities are demonstrated for an autonomous additive manufacturing process with a real-time feedback loop.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#engelmann26building\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>S. Boehm, C. A. Bridges, P. Widener, T. Jones, S. Ghafoor, C. Engelmann, and O. Kuchar. <b>The INTERSECT Scientific Data Layer: An Ontological Framework for Data Provenance for Complex Scientific Workflows<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/2026.euro-par.org\" target=\"2026.euro-par.org\">32nd European Conference on Parallel and Distributed Computing (Euro-Par) 2026 Workshops<\/a>: <a href=\"http:\/\/www.hipes-workshop.org\/\" target=\"www.hipes-workshop.org\/\">3rd Workshop on High-Performance eScience Tools and Applications (HiPES)<\/a><\/i>, August, 2026. To appear. <a href=\"javascript:showAbstract('Complex scientific workflows have multiple stages, including experiments, simulations, data analyses, and visualization, generating data and metadata stored across heterogeneous storage infrastructures and used in downstream stages or future experimental campaigns. This paper presents the design and implementation of a comprehensive ontological framework for managing scientific data generated within such workflows, addressing the key challenges of data interoperability, provenance capture, and adherence to Findable, Accessible, Interoperable, and Reusable (FAIR) data principles. The Autonomous Chemistry Laboratory (ACL) at Oak Ridge National Laboratory (ORNL) enables automated liquid phase and solid state synthesis and related chemical analysis. Our framework has been deployed within the ACL for a native and machine-interpretable semantic representation of its ecosystem, including instrument capabilities, synthesis workflows, analytical observations, and experimental results. The proposed approach enables end-to-end provenance tracking, supports heterogeneous data formats, and establishes a foundation for Artificial Intelligence (AI)-ready scientific discovery. We validate the framework through concrete modeling examples.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"?page_id=55#boehm28intersect\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>O. Kotevska, T. Nguyen, R. F. da Silva, C. Engelmann, and P. Balaprakash. <b>Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/2026.eurosys.org\/\" target=\"2026.eurosys.org\/\">21st European Conference on Computer Systems (EuroSyS)<\/a>: <a href=\"http:\/\/euromlsys.eu\/\" target=\"euromlsys.eu\/\">6th European Workshop on Machine Learning and Systems (EuroMLSys)<\/a><\/i>, April, 2026. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3805621.3807639\" target=\"publication\">10.1145\/3805621.3807639<\/a>. Accept. rate 69.2% (18\/26). <a href=\"javascript:showAbstract('Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronization and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kotevska26scalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kotevska26scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>P. Valero-Lara, A. Young, T. Naughton, C. Engelmann, A. Geist, J. S. Vetter, K. Teranishi, and W. F. Godoy. <b>ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads<\/b>. In <i>Proceedings of the <a href=\"http:\/\/www.sca-hpcasia2026.jp\" target=\"www.sca-hpcasia2026.jp\">Supercomputing Asia \/ International Conference on High Performance Computing in the Asia-Pacific Region (SCA\/HPCAsia) 2026<\/a><\/i>, January, 2026. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3773656.3773659\" target=\"publication\">10.1145\/3773656.3773659<\/a>. Accept. rate 36.6% (37\/101). <a href=\"javascript:showAbstract('The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually--especially applying a proper domain decomposition and communication pattern--is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)-based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4x boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/valero-lara26chatmpi.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#valero-lara26chatmpi\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>K. Kim, K. Raghavan, O. Kotevska, M. Dorier, R. Madduri, M. Ryu, T. Munson, R. Ross, T. Flynn, A. Kagawa, B. Yoon, C. Engelmann, and F. Yousefian. <b>Privacy-Preserving Federated Learning for Science: Challenges and Research Directions<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ieeebigdata2024.github.io\" target=\"ieeebigdata2024.github.io\">12th IEEE International Conference on Big Data  (BigData) 2024<\/a><\/i>, December, 2024. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/BigData62323.2024.10825853\" target=\"publication\">10.1109\/BigData62323.2024.10825853<\/a>. Accept. rate 18.5% (122\/661). <a href=\"javascript:showAbstract('This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kim24privacy.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kim24privacy\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Highly Cited Peer-Reviewed Publications<\/h4>\n<ol>\n<li>M. Snir, R. W. Wisniewski, J. A. Abraham, S. V. Adve, S. Bagchi, P. Balaji, J. Belak, P. Bose, F. Cappello, B. Carlson, A. A. Chien, P. Coteus, N. A. Debardeleben, P. Diniz, C. Engelmann, M. Erez, S. Fazzari, A. Geist, R. Gupta, F. Johnson, S. Krishnamoorthy, S. Leyffer, D. Liberty, S. Mitra, T. Munson, R. Schreiber, J. Stearley, and E. V. Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i><a href=\"http:\/\/hpc.sagepub.com\" target=\"hpc.sagepub.com\">International Journal of High Performance Computing Applications (IJHPCA)<\/a><\/i>, volume 28, number 2, May, 2014. DOI <a href=\"http:\/\/dx.doi.org\/10.1177\/1094342014522573\" target=\"publication\">10.1177\/1094342014522573<\/a>. 555 citations. <a href=\"javascript:showAbstract('We present here a report produced by a workshop on  Addressing failures in exascale computing&amp;#39; held in Park City,  Utah, 4-11 August 2012. The charter of this workshop was to  establish a common taxonomy about resilience across all the  levels in a computing system, discuss existing knowledge on  resilience across the various hardware and software layers  of an exascale system, and build on those results, examining  potential solutions from both a hardware and software  perspective and focusing on a combined approach. The workshop brought together participants with expertise in  applications, system software, and hardware; they came from  industry, government, and academia, and their interests ranged  from theory to implementation. The combination allowed broad  and comprehensive discussions and led to this document, which  summarizes and builds on those discussions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/snir14addressing.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir14addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>A. B. Nagarajan, F. Mueller, C. Engelmann, and S. L. Scott. <b>Proactive Fault Tolerance for HPC with Xen Virtualization<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ics07.ac.upc.edu\" target=\"ics07.ac.upc.edu\">21st ACM International Conference on Supercomputing (ICS) 2007<\/a><\/i>, June, 2007. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1274971.1274978\" target=\"publication\">10.1145\/1274971.1274978<\/a>. Accept. rate 23.6% (29\/123). 526 citations. <a href=\"javascript:showAbstract('Large-scale parallel computing is relying increasingly on clusters with thousands of processors. At such large counts of compute nodes, faults are becoming common place. Current techniques to tolerate faults focus on reactive schemes to recover from faults and generally rely on a checkpoint\/restart mechanism. Yet, in today`s systems, node failures can often be anticipated by detecting a deteriorating health status. Instead of a reactive scheme for fault tolerance (FT), we are promoting a proactive one where processes automatically migrate from unhealthy nodes to healthy ones. Our approach relies on operating system virtualization techniques exemplified by but not limited to Xen. This paper contributes an automatic and transparent mechanism for proactive FT for arbitrary MPI applications. It leverages virtualization techniques combined with health monitoring and load-based migration. We exploit Xen`s live migration mechanism for a guest operating system (OS) to migrate an MPI task from a health-deteriorating node to a healthy one without stopping the MPI task during most of the migration. Our proactive FT daemon orchestrates the tasks of health monitoring, load determination and initiation of guest OS migration. Experimental results demonstrate that live migration hides migration costs and limits the overhead to only a few seconds making it an attractive approach to realize FT in HPC systems. Overall, our enhancements make proactive FT a valuable asset for long-running MPI application that is complementary to reactive FT using full checkpoint\/restart schemes since checkpoint frequencies can be reduced as fewer unanticipated failures are encountered. In the context of OS virtualization, we believe that this is the first comprehensive study of proactive fault tolerance where live migration is actually triggered by health monitoring.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/nagarajan07proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/nagarajan07proactive.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#nagarajan07proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>D. Fiala, F. Mueller, C. Engelmann, K. Ferreira, R. Brightwell, and R. Riesen. <b>Detection and Correction of Silent Data Corruption for Large-Scale High-Performance Computing<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc12.supercomputing.org\" target=\"sc12.supercomputing.org\">25th IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2012<\/a><\/i>, November, 2012. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SC.2012.49\" target=\"publication\">10.1109\/SC.2012.49<\/a>. Accept. rate 21.2% (100\/472). 397 citations. <a href=\"javascript:showAbstract('Faults have become the norm rather than the exception for high-end computing on clusters with 10s\/100s of thousands of cores. Exacerbating this situation, some of these faults remain undetected, manifesting themselves as silent errors that corrupt memory while applications continue to operate and report incorrect results. This paper studies the potential for redundancy to both detect and correct soft errors in MPI message-passing applications. Our study investigates the challenges inherent to detecting soft errors within MPI application while providing transparent MPI redundancy. By assuming a model wherein corruption in application data manifests itself by producing differing MPI message data between replicas, we study the best suited protocols for detecting and correcting MPI data that is the result of corruption. To experimentally validate our proposed detection and correction protocols, we introduce RedMPI, an MPI library which resides in the MPI profiling layer. RedMPI is capable of both online detection and correction of soft errors that occur in MPI applications without requiring any modifications to the application source by utilizing either double or triple redundancy. Our results indicate that our most efficient consistency protocol can successfully protect applications experiencing even high rates of silent data corruption with runtime overheads between 0% and 30% as compared to unprotected applications without redundancy. Using our fault injector within RedMPI, we observe that even a single soft error can have profound effects on running applications, causing a cascading pattern of corruption in most cases causes that spreads to all other processes. RedMPI&amp;#39;s protection has been shown to successfully mitigate the effects of soft errors while allowing applications to complete with correct results even in the face of errors.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala12detection2.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/fiala12detection2.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#fiala12detection2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>C. Wang, F. Mueller, C. Engelmann, and S. L. Scott. <b>Proactive Process-Level Live Migration in HPC Environments<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc08.supercomputing.org\" target=\"sc08.supercomputing.org\">21st IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2008<\/a><\/i>, November, 2008. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/1413370.1413414\" target=\"publication\">10.1145\/1413370.1413414<\/a>. Accept. rate 21.3% (59\/277). 249 citations. <a href=\"javascript:showAbstract('As the number of nodes in high-performance computing environments keeps increasing, faults are becoming common place. Reactive fault tolerance (FT) often does not scale due to massive I\/O requirements and relies on manual job resubmission. This work complements reactive with proactive FT at the process level. Through health monitoring, a subset of node failures can be anticipated when one&amp;#39;s health deteriorates. A novel process-level live migration mechanism supports continued execution of applications during much of processes migration. This scheme is integrated into an MPI execution environment to transparently sustain health-inflicted node failures, which eradicates the need to restart and requeue MPI jobs. Experiments indicate that 1-6.5 seconds of prior warning are required to successfully trigger live process migration while similar operating system virtualization mechanisms require 13-24 seconds. This self-healing approach complements reactive FT by nearly cutting the number of checkpoints in half when 70% of the faults are handled proactively.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/wang08proactive.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/wang08proactive.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#wang08proactive\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>J. Elliott, K. Kharbas, D. Fiala, F. Mueller, K. Ferreira, and C. Engelmann. <b>Combining Partial Redundancy and Checkpointing for HPC<\/b>. In <i>Proceedings of the <a href=\"http:\/\/icdcs-2012.org\/\" target=\"icdcs-2012.org\/\">32nd International Conference on Distributed Computing Systems (ICDCS) 2012<\/a><\/i>, June, 2012. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/ICDCS.2012.56\" target=\"publication\">10.1109\/ICDCS.2012.56<\/a>. Accept. rate 13.8% (71\/515). 211 citations. <a href=\"javascript:showAbstract('Today&amp;#39;s largest High Performance Computing (HPC) systems exceed one Petaflops (10^15 floating point operations per second) and exascale systems are projected within seven years. But reliability is becoming one of the major challenges faced by exascale computing. With billion-core parallelism, the mean time to failure is projected to be in the range of minutes or hours instead of days. Failures are becoming the norm rather than the exception during execution of HPC applications. Current fault tolerance techniques in HPC focus on reactive ways to mitigate faults, namely via checkpoint and restart (C\/R). Apart from storage overheads, C\/R-based fault recovery comes at an additional cost in terms of application performance because normal execution is disrupted when checkpoints are taken. Studies have shown that applications running at a large scale spend more than 50% of their total time saving checkpoints, restarting and redoing lost work. Redundancy is another fault tolerance technique, which employs redundant processes performing the same task. If a process fails, a replica of it can take over its execution. Thus, redundant copies can decrease the overall failure rate. The downside of redundancy is that extra resources are required and there is an additional overhead on communication and synchronization. This work contributes a model and analyzes the benefit of C\/R in coordination with redundancy at different degrees to minimize the total wallclock time and resources utilization of HPC applications. We further conduct experiments with an implementation of redundancy within the MPI layer on a cluster. Our experimental results confirm the benefit of dual and triple redundancy - but not for partial redundancy - and show a close fit to the model. At 80,000 processes, dual redundancy requires twice the number of processing resources for an application but allows two jobs of 128 hours wallclock time to finish within the time of just one job without redundancy. For narrow ranges of processor counts, partial redundancy results in the lowest time. Once the count exceeds 770, 000, triple redundancy has the lowest overall cost. Thus, redundancy allows one to trade-off additional resource requirements against wallclock time, which provides a tuning knob for users to adapt to resource availabilities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/elliott12combining.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/elliott12combining.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#elliott12combining\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Other Significant Publications<\/h4>\n<ol>\n<li>M. Kumar, S. Gupta, T. Patel, M. Wilder, W. Shi, S. Fu, C. Engelmann, and D. Tiwari. <b>Study of Interconnect Errors, Network Congestion, and Applications Characteristics for Throttle Prediction on a Large Scale HPC System<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/jpdc\" target=\"www.elsevier.com\/locate\/jpdc\">Journal of Parallel and Distributed Computing (JPDC)<\/a><\/i>, volume 153, July, 2021. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.jpdc.2021.03.001\" target=\"publication\">10.1016\/j.jpdc.2021.03.001<\/a>. <a href=\"javascript:showAbstract('Today&amp;#39;s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of peta ops. The performance of these systems depends on various components like processors, memory, and interconnect. Among  all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kumar21study.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kumar21study\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>G. Ostrouchov, D. Maxwell, R. Ashraf, C. Engelmann, M. Shankar, and J. Rogers. <b>GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability<\/b>. In <i>Proceedings of the <a href=\"http:\/\/sc20.supercomputing.org\" target=\"sc20.supercomputing.org\">33rd IEEE\/ACM International Conference on High Performance Computing, Networking, Storage and Analysis (SC) 2020<\/a><\/i>, November, 2020. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/SC41405.2020.00045\" target=\"publication\">10.1109\/SC41405.2020.00045<\/a>. Accept. rate 25.1% (95\/378). <a href=\"javascript:showAbstract('The Cray XK7 Titan was the top supercomputer system in the world for a very long time and remained critically important throughout its nearly seven year life. It was also a very interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three very significant rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 operation years in the GPU lifetimes, which correspond to Titan&amp;#39;s 6 year long productive period after an initial break-in period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the system cooling architecture and job scheduling. In addition to describing some of the system history, the data collection, data cleaning, and our analysis of the data, we provide reliability recommendations for designing future state of the art supercomputing systems and their operation. We make the data and our analysis codes publicly available.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/ostrouchov20gpu.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/ostrouchov20gpu.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#ostrouchov20gpu\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>H. Jeong, Y. Yang, C. Engelmann, V. Gupta, T. M. Low, P. Grover, V. Cadambe, and K. Ramchandran. <b>3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication<\/b>. In <i>Lecture Notes in Computer Science: Proceedings of the <a href=\"http:\/\/www.euro-par.org\" target=\"www.euro-par.org\">26th European Conference on Parallel and Distributed Computing (Euro-Par) 2020<\/a><\/i>, August, 2020. DOI <a href=\"http:\/\/dx.doi.org\/10.1007\/978-3-030-57675-2_25\" target=\"publication\">10.1007\/978-3-030-57675-2_25<\/a>. Accept. rate 24.5% (39\/159). <a href=\"javascript:showAbstract('In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that is communication efficient and achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for distributed matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires 50% less redundancy than replication, while the overhead in execution time is only about 5-10%.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/jeong203d.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/jeong203d.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#jeong203d\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>D. Fiala, F. Mueller, K. Ferreira, and C. Engelmann. <b>Mini-Ckpts: Surviving OS Failures in Persistent Memory<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ics16.bilkent.edu.tr\" target=\"ics16.bilkent.edu.tr\">30th ACM International Conference on Supercomputing  (ICS) 2016<\/a><\/i>, June, 2016. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/2925426.2926295\" target=\"publication\">10.1145\/2925426.2926295<\/a>. Accept. rate 24.2% (43\/178). <a href=\"javascript:showAbstract('Concern is growing in the high-performance computing (HPC) community on the reliability of future extreme-scale systems. Current efforts have focused on application fault-tolerance rather than the operating system (OS), despite the fact that recent studies have suggested that failures in OS memory are more likely. The OS is critical to a system&amp;#39;s correct and efficient operation of the node and processes it governs -- and in HPC also for any other nodes a parallelized application runs on and communicates with: Any single node failure generally forces all processes of this application to terminate due to tight communication in HPC. Therefore, the OS itself must be capable of tolerating failures. In this work, we introduce mini-ckpts, a framework which enables application survival despite the occurrence of a fatal OS failure or crash. Mini-ckpts achieves this tolerance by ensuring that the critical data describing a process is preserved in persistent memory prior to the failure. Following the failure, the OS is rejuvenated via a warm reboot and the application continues execution effectively making the failure and restart transparent. The mini-ckpts rejuvenation and recovery process is measured to take between three to six seconds and has a failure-free overhead of between 3-5% for a number of key HPC workloads. In contrast to current fault-tolerance methods, this work ensures that the operating and runtime system can continue in the presence of faults. This is a much finer-grained and dynamic method of fault-tolerance than the current, coarse-grained, application-centric methods. Handling faults at this level has the potential to greatly reduce overheads and enables mitigation of additional fault scenarios.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/fiala16mini-ckpts.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/fiala16mini-ckpts.ppt.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#fiala16mini-ckpts\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>C. Engelmann. <b>Scaling To A Million Cores And Beyond: Using Light-Weight Simulation to Understand The Challenges Ahead On The Road To Exascale<\/b>. <i><a href=\"http:\/\/www.elsevier.com\/locate\/fgcs\" target=\"www.elsevier.com\/locate\/fgcs\">Future Generation Computer Systems (FGCS)<\/a><\/i>, volume 30, number 0, January, 2014. DOI <a href=\"http:\/\/dx.doi.org\/10.1016\/j.future.2013.04.014\" target=\"publication\">10.1016\/j.future.2013.04.014<\/a>. 69 citations. <a href=\"javascript:showAbstract('As supercomputers scale to 1,000 PFlop\/s over the next decade, investigating the performance of parallel applications at scale on future architectures and the performance impact of different architecture choices for high-performance computing (HPC) hardware\/software co-design is crucial. This paper summarizes recent efforts in designing and implementing a novel HPC hardware\/software co-design toolkit. The presented Extreme-scale Simulator (xSim) permits running an HPC application in a controlled environment with millions of concurrent execution threads while observing its performance in a simulated extreme-scale HPC system using architectural models and virtual timing. This paper demonstrates the capabilities and usefulness of the xSim performance investigation toolkit, such as its scalability to 2^27 simulated Message Passing Interface (MPI) ranks on 960 real processor cores, the capability to evaluate the performance of different MPI collective communication algorithms, and the ability to evaluate the performance of a basic Monte Carlo application with different architectural parameters.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13scaling.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann13scaling\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Extreme-Scale Computing | Fault Resilience | HW\/SW Co-Design Tools | Computing Continuum | Autonomous Experiments Dr. Christian Engelmann is a Distinguished Computer Scientist and the Intelligent Systems and Facilities Research Group Leader at Oak Ridge National Laboratory (ORNL), the US Department of Energy&#8217;s (DOE) largest multiprogram science and technology laboratory with an annual budget of&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-2","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/2","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2"}],"version-history":[{"count":300,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/2\/revisions"}],"predecessor-version":[{"id":1455,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/2\/revisions\/1455"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}