{"id":45,"date":"2023-02-10T08:00:11","date_gmt":"2023-02-10T08:00:11","guid":{"rendered":"https:\/\/christian-engelmann.de\/?page_id=45"},"modified":"2023-02-15T20:19:17","modified_gmt":"2023-02-15T20:19:17","slug":"whitepapers","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=45","title":{"rendered":"Whitepapers"},"content":{"rendered":"<ol>\n<li>Ryan Adamson and Christian Engelmann. <b>Cybersecurity and Privacy for Instrument-to-Edge-to-Center Scientific Computing Ecosystems<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/2021ascr-cybersecurity\" target=\"www.orau.gov\/2021ascr-cybersecurity\" rel=\"noopener\">ASCR Workshop on Cybersecurity and Privacy for Scientific  Computing Ecosystems<\/a><\/i>, November 3-5, 2021. <a href=\"javascript:showAbstract('The DOE&amp;#39;s Artificial Intelligence (AI) for Science report outlines the need for intelligent systems, instruments, and facilities to enable science breakthroughs with autonomous experiments, 'self-driving' laboratories, smart manufacturing, and AI-driven design, discovery and evaluation. The DOE's Computational Facilities Research Workshop report identifies intelligent systems\/facilities as a challenge with enabling automation and eliminating human-in-the-loop needs as a cross-cutting theme. Autonomous experiments, 'self-driving' laboratories and smart manufacturing employ machine-in-the-loop intelligence for decision-making. Human-in-the-loop needs are reduced by an autonomous online control that collects experiment data, analyzes it, and takes appropriate operational actions in real time to steer an ongoing or plan the next experiment. DOE laboratories are currently in the process of developing and deploying federated hardware\/software architectures for connecting instruments with edge and center computing resources to autonomously collect, transfer, store, process, curate, and archive scientific data. These new instrument-to-edge-to-center scientific ecosystems face several cybersecurity and privacy challenges.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/adamson21cybersecurity.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#adamson21cybersecurity\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Mingyan Li, Robert A. Bridges, Pablo Moriano, Christian Engelmann, Feiyi Wang, and Ryan Adamson. <b>Toward Effective Security\/Reliability Situational Awareness via Concurrent Security-or-Fault Analytics <\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/2021ascr-cybersecurity\" target=\"www.orau.gov\/2021ascr-cybersecurity\" rel=\"noopener\">ASCR Workshop on Cybersecurity and Privacy for Scientific  Computing Ecosystems<\/a><\/i>, November 3-5, 2021. <a href=\"javascript:showAbstract('Modern critical infrastructures (CI) and scientific computing ecosystems (SCE) are complex and vulnerable. The complexity of CI\/SCE, such as the distributed workload found across ASCR scientific computing facilities, does not allow for easy differentiation between emerging cyber security and reliability threats. It is also not easy to correctly identify the misbehaving systems. Sometimes, system failures are just caused by unintentional user misbehavior or actual hardware\/software reliability issues, but it may take some significant amount of time and effort to develop that understanding through root-cause analysis. On the security front, CI\/SCE are vital assets. They are prime targets of, and are vulnerable to, malicious cyber-attacks. Within DoE, inter-disciplinary and cross-facility collaboration (e.g., ORNL INTERSECT initiative, next-gen supercomputing OLCF6), traditional perimeter-based defense and demarcation line between malicious cyber-attacks and non-malicious system faults are blurring. Amidst realistic reliability and security threats, the ability to effectively distinguish between non-malicious faults and malicious attacks is critical not only in root cause identification but also in countermeasures generation. ');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/li21toward.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#li21toward\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Hal Finkel, Pete Beckman, Christian Engelmann, Shantenu Jha, and Jack Lange. <b>Research Opportunities in Operating Systems for Scientific Edge Computing<\/b>. <i>White paper by the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/OSRoundtable2021\" target=\"www.orau.gov\/OSRoundtable2021\" rel=\"noopener\">ASCR Roundtable Discussions on Operating-Systems Research 2021<\/a><\/i>, January 25, 2021. <a href=\"javascript:showAbstract('As scientific experiments generate ever-increasing amounts of data, and grow in operational complexity, modern experimental science demands unprecedented computational capabilities at the edge - physically proximate to each experiment. While some requirements on these computational capabilities are shared with high-performance-computing (HPC) systems, scientific edge computing has a number of unique challenges. In the following, we survey current trends in system software and edge systems for scientific computing, associated research challenges and open questions, infrastructure requirements for operating-systems research, communities who should be involved in that research, and the anticipated benefits of success.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/finkel21research2.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#finkel21research2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Hal Finkel, Pete Beckman, Ron Brightwell, Rudi Eigenmann, Christian Engelmann, Roberto Gioiosa, Kamil Iskra, Shantenu Jha, Jack Lange, Tapasya Patki, and Kevin Pedretti. <b>Research Opportunities in Operating Systems for High-Performance Scientific Computing<\/b>. <i>White paper by the U.S. Department of Energy&#39;s <a href=\"http:\/\/www.orau.gov\/OSRoundtable2021\" target=\"www.orau.gov\/OSRoundtable2021\" rel=\"noopener\">ASCR Roundtable Discussions on Operating-Systems Research 2021<\/a><\/i>, January 25, 2021. <a href=\"javascript:showAbstract('As high-performance-computing (HPC) systems continue to evolve, with increasingly diverse and heterogeneous hardware, increasingly-complex requirements for security and multi-tenancy, and increasingly-demanding requirements for resiliency and monitoring, research in operating systems must continue to seed innovation to meet future needs. In the following, we survey current trends in system software and HPC systems for scientific computing, associated research challenges and open questions, infrastructure requirements for operating-systems research, communities who should be involved in that research, and the anticipated benefits of success.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/finkel21research.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#finkel21research\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann. <b>Resilience by Codesign (and not as an Afterthought)<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/web.cvent.com\/event\/f64a4f28-b473-4808-924c-c8c3d9a2af63\/\" target=\"web.cvent.com\/event\/f64a4f28-b473-4808-924c-c8c3d9a2af63\/\" rel=\"noopener\">Workshop on Reimagining Codesign 2021<\/a><\/i>, March 16-18, 2021. <a href=\"javascript:showAbstract('Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale high-performance computing (HPC). Extreme heterogeneity, i.e., using multiple, and potentially configurable, types of processors, accelerators and memory\/storage in a single computing platform, will add a significant amount of complexity to the HPC hardware\/software eco-system. Hardware\/software HPC codesign for resilience is mostly nonexistent at this point! Resilience needs to become an integral part of the HPC hardware\/software ecosystem through codesign, such that the burden for resilience is on the system by design and not on the operator or user as an afterthought. Simply put, if resilience by design is not done now, in the early stages of extreme heterogeneity, the current state of practice for HPC resilience, global application-level checkpoint\/restart, will re-main the same for decades to come due to the high costs of adoption of alternatives later on. ');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann21resilience2.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann21resilience2.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann21resilience2\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Petar Radojkovic, Manolis Marazakis, Paul Carpenter, Reiley Jeyapaul, Dimitris Gizopoulos, Martin Schulz, Adria Armejach, Eduard Ayguade, Fran&ccedil;ois Bodin, Ramon Canal, Franck Cappello, Fabien Chaix, Guillaume Colin de Verdiere, Said Derradji, Stefano Di Carlo, Christian Engelmann, Ignacio Laguna, Miquel Moreto, Onur Mutlu, Lazaros Papadopoulos, Olly Perks, Manolis Ploumidis, Bezhad Salami, Yanos Sazeides, Dimitrios Soudris, Yiannis Sourdis, Per Stenstrom, Samuel Thibault, Will Toms, and Osman Unsal. <b>Towards Resilient EU HPC Systems: A Blueprint<\/b>. <i>White paper by the <a href=\"http:\/\/resilienthpc.eu\" target=\"resilienthpc.eu\" rel=\"noopener\">European HPC resilience initiative<\/a><\/i>, April 9, 2020. <a href=\"javascript:showAbstract('This document aims to spearhead a Europe-wide discussion on HPC system resilience and to help the European HPC community define best practices for resilience. We analyse a wide range of state-of-the-art resilience mechanisms and recommend the most effective approaches to employ in large-scale HPC systems. Our guidelines will be useful in the allocation of available resources, as well as guiding researchers and research funding towards the enhancement of resilience approaches with the highest priority and utility. Although our work is focussed on the needs of next generation HPC systems in Europe, the principles and evaluations are applicable globally. This document is the first output of the ongoing European HPC resilience initiative and it covers individual nodes in HPC systems, encompassing CPU, memory, intra-node interconnect and emerging FPGA-based hardware accelerators. With community support and feedback on this initial document, we will update the analysis and expand the scope to include other types of accelerators, as well as networks and storage.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/radojkovic20towards.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#radojkovic20towards\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Rizwan Ashraf, and Saurabh Hukerikar. <b>Extreme Heterogeneity with Resilience by Design (and not as an Afterthought)<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/orau.gov\/exheterogeneity2018\/\" target=\"orau.gov\/exheterogeneity2018\/\" rel=\"noopener\">Extreme Heterogeneity Virtual Workshop 2018<\/a><\/i>, January 23-24, 2018. <a href=\"javascript:showAbstract('Resilience, i.e., obtaining a correct solution in a timely and efficient manner, is one of the key challenges in extreme-scale high-performance computing (HPC). Extreme heterogeneity, i.e., using multiple, and potentially configurable, types of processors, accelerators and  memory\/storage in a single computing platform, will add a significant amount of complexity to the HPC hardware\/software ecosystem. The notion of correct computation and program state assumed by users and application developers today, which has been based on binary bit-level correctness, will no longer hold for processing elements based on quantum qubits and analog circuits that model spiking neurons in neuromorphic computing elements. The diverse set of compute and memory components in future heterogeneous systems will require novel hardware and software resilience solutions. Errors and failures reported by such heterogeneous hardware will need to be handled by the appropriate software component to enable efficient masking, recovery, and avoidance with little burden on the user. Similarly, errors and failures reported by the software running on such heterogeneous hardware need to be equally efficiently handled with little burden on the user. This requires a new approach, where resilience is holistically provided by the HPC hardware\/software ecosystem. The key challenges are to design and to operate extreme heterogeneous HPC systems with (1) wide-ranging resilience capabilities in system software, programming models, libraries, and applications, (2) interfaces and mechanisms for coordinating resilience capabilities across diverse hardware and software components, (3) appropriate metrics and tools for assessing performance, resilience, and energy, and (4) an understanding of the performance, resilience and energy trade-off that eventually results in well-informed HPC system design choices and runtime decisions.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann18extreme.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann18extreme\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Devesh Tiwari, Saurabh Gupta, and Christian Engelmann. <b>Lightweight, Actionable Analytical Tools Based on Statistical Learning for Efficient System Operations<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2016\" target=\"hpc.pnl.gov\/modsim\/2016\" rel=\"noopener\">Workshop on Modeling &#038; Simulation of Systems &#038; Applications (ModSim) 2016<\/a><\/i>, August 10-12, 2016. <a href=\"javascript:showAbstract('Modeling and simulation community has always relied on accurate and meaningful system data and parameters to drive analytical models and simulators. HPC systems continuously generate huge amount system event related data (e.g., system log, resource consumption log, RAS logs, power consumption logs), but meaningful interpretation and accuracy verification of such data is quite challenging. This talk offers a unique perspective and experience in demonstrating how modeling and simulation based research can actually be translated into production systems. We will discuss the short-term opportunities for modeling and simulation community to increase the impact and effectiveness of our analytical tools, &amp;#34;dos and don&amp;#39;ts&amp;#34;, long-term challenges and opportunities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/tiwari16lightweight.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/tiwari16lightweight.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#tiwari16lightweight\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A Hardware\/Software Performance\/Resilience\/Power Co-Design Tool for Extreme-scale Computing<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2013\" target=\"hpc.pnl.gov\/modsim\/2013\" rel=\"noopener\">Workshop on Modeling &#038; Simulation of Exascale Systems &#038; Applications (ModSim) 2013<\/a><\/i>, September 18-19, 2013. <a href=\"javascript:showAbstract('xSim is a simulation-based performance investigation toolkit that permits running high-performance computing (HPC) applications in a controlled environment with millions of concurrent execution threads, while observing application performance in a simulated extreme-scale system for hardware\/software co-design. The presented work details newly developed features for xSim that permit the injection of MPI process failures, the propagation\/detection\/notification of such failures within the simulation, and their handling using application-level checkpoint\/restart. The newly added features also offer user-level failure mitigation (ULFM) extensions at the simulated MPI layer to support algorithm-based fault tolerance (ABFT). The presented solution permits investigating performance under failure and failure handling of checkpoint\/restart and ABFT solutions. The newly enhanced xSim is the very first performance tool that supports these capabilities.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann13hardware.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann13hardware.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann13hardware\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Marc Snir, and Robert W. Wisniewski, Jacob A. Abraham, Sarita V. Adve, Saurabh Bagchi, Pavan Balaji, Bill Carlson, Andrew A. Chien, Pedro Diniz, Christian Engelmann, Rinku Gupta, Fred Johnson, Jim Belak, Pradip Bose, Franck Cappello, Paul Coteus, Nathan A. Debardeleben, Mattan Erez, Saverio Fazzari, Al Geist, Sriram Krishnamoorthy, Sven Leyffer, Dean Liberty, Subhasish Mitra, Todd Munson, Rob Schreiber, Jon Stearley, and Eric Van Hensbergen. <b>Addressing Failures in Exascale Computing<\/b>. <i>Workshop report<\/i>, August 4-11, 2013. <a href=\"publications\/snir13addressing.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#snir13addressing\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Al Geist, Bob Lucas, Marc Snir, Shekhar Borkar, Eric Roman, Mootaz Elnozahy, Bert Still, Andrew Chien, Robert Clay, John Wu, Christian Engelmann, Nathan DeBardeleben, Rob Ross, Larry Kaplan, Martin Schulz, Mike Heroux, Sriram Krishnamoorthy, Lucy Nowell, Abhinav Vishnu, and Lee-Ann Talley. <b>U.S. Department of Energy Fault Management Workshop<\/b>. <i>Workshop report for the U.S. Department of Energy<\/i>, June 6, 2012. <a href=\"javascript:showAbstract('A Department of Energy (DOE) Fault Management Workshop was held on June 6, 2012 at the BWI Airport Marriot hotel in Maryland. The goals of this workshop were to: 1. Describe the required HPC resilience for critical DOE mission needs; 2. Detail what HPC resilience research is already being done at the DOE national laboratories and is expected to be done by industry or other groups; 3. Determine what fault management research is a priority for DOE&amp;#39;s Office of Science and National Nuclear Security Administration (NNSA) over the next five years; 4. Develop a roadmap for getting the necessary research accomplished in the timeframe when it will be needed by the large computing facilities across DOE.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/geist12department.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#geist12department\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann and Thomas Naughton. <b>A Performance\/Resilience\/Power Co-design Tool for Extreme-scale High-Performance Computing<\/b>. <i>White paper accepted at the U.S. Department of Energy&#39;s <a href=\"http:\/\/hpc.pnl.gov\/modsim\/2012\" target=\"hpc.pnl.gov\/modsim\/2012\" rel=\"noopener\">Workshop on Modeling &#038; Simulation of Exascale Systems &#038; Applications (ModSim) 2012<\/a><\/i>, August 9-10, 2012. <a href=\"javascript:showAbstract('Performance, resilience and power consumption are key HPC system design factors that are highly interde-pendent. To enable extreme-scale computing it is essential to perform HPC hardware\/software co-design that identifies the cost\/benefit trade-off between these design factors for potential future architecture choices. The proposed research and development aims at developing an HPC hardware\/software co-design toolkit for evaluating the resilience\/power\/performance cost\/benefit trade-off of future architecture choices. The approach focuses on extending a simulation-based performance investigation toolkit with advanced resilience and power modeling and simulation features, such as (i) fault injection mechanisms, (ii) fault propagation, isolation, and detection models, (i) fault avoidance, masking, and recovery simulation, and (iv) power consumption models.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12performance.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#engelmann12performance\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Christian Engelmann, Geoffroy R. Vall&eacute;e, Thomas Naughton, and Frank Mueller. <b>Dynamic Self-Aware Runtime Software for Exascale Systems<\/b>. <i>White paper for the U.S. Department of Energy&#39;s <a href=\"http:\/\/collab.cels.anl.gov\/display\/exaosr\/Position+Papers\" target=\"collab.cels.anl.gov\/display\/exaosr\/Position+Papers\" rel=\"noopener\">Exascale Operating Systems and Runtime Technical Council<\/a><\/i>, July 1, 2012. <a href=\"javascript:showAbstract('At exascale, the power consumption, resilience, and load balancing constraints, especially their dynamic nature and interdependence, and the scale of the system require a radical change in future high-performance computing (HPC) operating systems and runtimes (OS\/Rs). In contrast to the existing static OS\/R solutions, an exascale OS\/R is needed that is aware of the dynamically changing resources, constraints, and application needs, and that is able to autonomously coordinate (sometimes conflicting) responses to different changes in the system, simultaneously and at scale. To provide awareness and autonomic management, a novel, scalable and self-aware OS\/R is needed that becomes the brains of the entire X-stack. It dynamically analyzes past, current, and future system status and application needs. It optimizes system usage by scheduling, migrating, and restarting tasks within and across nodes as needed to deal with multi-dimensional constraints, such as power consumption, permanent and transient faults, resource degradation, heterogeneity, data locality, and load balance.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/engelmann12dynamic.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"publications\/engelmann12dynamic.ppt.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/ppt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Presentation\"><\/a> <a href=\"?page_id=55#engelmann12dynamic\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Geoffroy R. Vall&eacute;e, Thomas Naughton, Christian Engelmann, and David E. Bernholdt. <b>Unified Execution Environment<\/b>. <i>White paper for the U.S. Department of Energy&#39;s <a href=\"http:\/\/collab.cels.anl.gov\/display\/exaosr\/Position+Papers\" target=\"collab.cels.anl.gov\/display\/exaosr\/Position+Papers\" rel=\"noopener\">Exascale Operating Systems and Runtime Technical Council<\/a><\/i>, July 1, 2012. <a href=\"publications\/vallee12unified.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#vallee12unified\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<li>Nathan DeBardeleben, James Laros, John T. Daly, Stephen L. Scott, Christian Engelmann, and Bill Harrod. <b>High-End Computing Resilience: Analysis of Issues Facing the HEC Community and Path-Forward for Research and Development<\/b>. <i>White paper for the U.S. National Science Foundation&#39;s High-end Computing Program<\/i>, December 1, 2009. <a href=\"publications\/debardeleben09high-end.pdf\" target=\"publication\" rel=\"noopener\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#debardeleben09high-end\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Ryan Adamson and Christian Engelmann. Cybersecurity and Privacy for Instrument-to-Edge-to-Center Scientific Computing Ecosystems. White paper accepted at the U.S. Department of Energy&#39;s ASCR Workshop on Cybersecurity and Privacy for Scientific Computing Ecosystems, November 3-5, 2021. Mingyan Li, Robert A. Bridges, Pablo Moriano, Christian Engelmann, Feiyi Wang, and Ryan Adamson. Toward Effective Security\/Reliability Situational Awareness via&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":16,"menu_order":4,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-45","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/45","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=45"}],"version-history":[{"count":3,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/45\/revisions"}],"predecessor-version":[{"id":191,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/45\/revisions\/191"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/16"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=45"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}