{"id":1188,"date":"2026-04-26T08:00:49","date_gmt":"2026-04-26T08:00:49","guid":{"rendered":"https:\/\/www.christian-engelmann.info\/?page_id=1188"},"modified":"2026-04-29T23:26:56","modified_gmt":"2026-04-29T23:26:56","slug":"2024-privacy-preserving-federated-learning-for-science-building-sustainable-and-trustworthy-foundation-models","status":"publish","type":"page","link":"https:\/\/www.christian-engelmann.info\/?page_id=1188","title":{"rendered":"2024-&#8230;: Privacy-Preserving Federated Learning for Science: Building Sustainable and Trustworthy Foundation Models"},"content":{"rendered":"<p>Federated learning (FL) offers a collaborative framework for training foundation models (FMs) and other AI models across distributed computing infrastructures and datasets while incorporating privacy-preserving techniques to manage private-sensitive datasets. This proposal addresses the challenges inherent in adapting FL to the &#8220;pre-train&#8221; and &#8220;fine-tune&#8221; paradigms of FMs with billions or trillions of parameters. These challenges include increased communication costs, computation burdens on clients, and the handling of massive model parameters and multi-modal datasets. Moreover, existing privacy-preserving techniques, such as differential privacy (DP), need to address scalability issues with such large models and different privacy requirements from clients. With synthetic data emerging as a promising alternative, new data management challenges are anticipated in the privacy-preserving FL (PPFL) framework with privacy-sensitive datasets and synthetic data.<\/p>\n<p>The project develops efficient communication, memory, and energy optimization techniques for FL algorithms, particularly for large-scale FMs, while ensuring fairness and incentivizing participation. It advances DP techniques to address scalability and heterogeneity challenges, create and manage synthetic data to preserve privacy while maintaining data utility, and integrate these efforts into a cohesive data management framework to enhance the scalability and performance of PPFL systems. Specifically, the research is structured around four main thrusts: (1) improving communication, memory, and energy efficiency; (2) addressing continual learning with incentives and fairness; (3) developing scalable and heterogeneous DP techniques; and (4) creating synthetic data as a privacy-preserving alternative. A crosscut thrust integrates these efforts, providing efficient model and data management schemes using tools, like Mofka and ProxyStore, to tackle access, sharing, versioning, control, and evolution of large datasets and models.<\/p>\n<p>This research effort significantly advances the field of PPFL by enhancing the scalability and efficiency of training large FMs, ensuring fairness and incentive structures for client participation in FL, developing scalable DP techniques that maintain model utility while ensuring privacy, and creating high-quality synthetic data as a proxy for sensitive datasets. The integration of these thrusts is demonstrated through specific scientific use cases in X-ray image science and electric grids, focusing on efficiently training large FMs with substantial data streams subject to privacy constraints. The outcomes ensure the sustainable and trustworthy training and deployment of FMs for science, benefiting a wide range of applications and advancing the state of the art in AI and FL. For more information, please visit <a href=\"https:\/\/ai4s-ppfl.github.io\" target=\"ai4s-ppfl.github.io\">ai4s-ppfl.github.io<\/a><\/p>\n<h4>Funding Sources<\/h4>\n<ul>\n<li>Advancements in Artificial Intelligence for Science Program, <a href=\"http:\/\/science.energy.gov\/ascr\" target=\"science.energy.gov_ascr\">Office of Advanced Scientific Computing Research<\/a>, Office of Science, U.S. Department of Energy\n  <\/li>\n<\/ul>\n<h4>Participants<\/h4>\n<ul>\n<li>Kibaek Kim (PI), Ravi Madduri, Todd Munson, Krishnan Raghavan, Rob Ross, and Matthieu Dorier &#8212; <a href=\"http:\/\/www.anl.gov\" target=\"www.anl.gov\">Argonne National Laboratory<\/a>\n<\/li>\n<li>Thomas Flynn, Ai Kagawa, and Byung-Jun Yoon &#8212; <a href=\"http:\/\/www.bnl.gov\" target=\"www.anl.gov\">Brookhaven National Laboratory<\/a>\n<\/li>\n<li>Olivera Kotevska and Christian Engelmann &#8212; <a href=\"http:\/\/www.ornl.gov\" target=\"www.ornl.gov\">Oak Ridge National Laboratory<\/a>\n<\/li>\n<li>Minseok Ryu &#8212; <a href=\"http:\/\/www.asu.edu\" target=\"www.ornl.gov\">Arizona State University<\/a>\n<\/li>\n<li>Farzad Yousefian &#8212; <a href=\"http:\/\/www.rutgers.edu\" target=\"www.ornl.gov\">Rutgers University<\/a>\n<\/li>\n<\/ul>\n<h4>In the News<\/h4>\n<p><b>2025-07-08:<\/b> DOE Advanced Scientific Computing Research. <a href=\"https:\/\/science.osti.gov\/-\/media\/ascr\/pdf\/facilities\/ALCC\/ALCC_Factsheets_2025.pdf\" target=\"ALCC_Factsheets_2025\">1.1 million supercomputer node-hours awarded to Privacy-Preserving Federated Learning for Foundation Models<\/a>.<br \/>\n<b>2024-10-15:<\/b> ORNL News. <a href=\"https:\/\/www.ornl.gov\/news\/ornl-projects-included-67-million-doe-ai-science-research\" target=\"www.ornl.gov_news_ornl-projects-included-67-million-doe-ai-science-research\">New ORNL projects included in $67 million from DOE for AI in science research<\/a>.\n<\/p>\n<h4>Peer-reviewed Conference Publications<\/h4>\n<ol>\n<li>Kibaek Kim, Krishnan Raghavan, Olivera Kotevska, Matthieu Dorier, Ravi Madduri, Minseok Ryu, Todd Munson, Rob Ross, Thomas Flynn, Ai Kagawa, Byung-Jun Yoon, Christian Engelmann, and Farzad Yousefian. <b>Privacy-Preserving Federated Learning for Science: Challenges and Research Directions<\/b>. In <i>Proceedings of the <a href=\"http:\/\/ieeebigdata2024.github.io\" target=\"ieeebigdata2024.github.io\">12th IEEE International Conference on Big Data  (BigData) 2024<\/a><\/i>, pages 7849-7853, Washington, DC, USA, December 15-18, 2024. <a href=\"http:\/\/www.computer.org\" target=\"www.computer.org\">IEEE Computer Society, Los Alamitos, CA, USA<\/a>. ISBN 979-8-3503-6249-7. ISSN 2639-1589. DOI <a href=\"http:\/\/dx.doi.org\/10.1109\/BigData62323.2024.10825853\" target=\"publication\">10.1109\/BigData62323.2024.10825853<\/a>. Acceptance rate 18.5% (122\/661). <a href=\"javascript:showAbstract('This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kim24privacy.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kim24privacy\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<h4>Peer-reviewed Workshop Publications<\/h4>\n<ol>\n<li>Olivera Kotevska, Trong Nguyen, Rafael Ferreira da Silva, Christian Engelmann, and Prasanna Balaprakash. <b>Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems<\/b>. In <i>Proceedings of the <a href=\"http:\/\/2026.eurosys.org\/\" target=\"2026.eurosys.org\/\">21st European Conference on Computer Systems (EuroSyS)<\/a>: <a href=\"http:\/\/euromlsys.eu\/\" target=\"euromlsys.eu\/\">6th European Workshop on Machine Learning and Systems (EuroMLSys)<\/a><\/i>, pages 439-446, Edinburgh, United Kingdom, April 27, 2026. <a href=\"http:\/\/www.acm.org\" target=\"www.acm.org\">ACM Press, New York, NY, USA<\/a>. ISBN 979-8-4007-2605-7. DOI <a href=\"http:\/\/dx.doi.org\/10.1145\/3805621.3807639\" target=\"publication\">10.1145\/3805621.3807639<\/a>. Acceptance rate 69.2% (18\/26). <a href=\"javascript:showAbstract('Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronization and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.');\"><img decoding=\"async\" src=\"images\/txt.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Abstract\"><\/a> <a href=\"publications\/kotevska26scalable.pdf\" target=\"publication\"><img decoding=\"async\" src=\"images\/pdf.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"Publication\"><\/a> <a href=\"?page_id=55#kotevska26scalable\"><img decoding=\"async\" src=\"images\/bib.gif\" border=\"0\" style=\"border-style:none\" height=\"10pt\" alt=\"BibTeX Citation\"><\/a><\/li>\n<\/ol>\n<p><em><small>Symbols: <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/txt.gif\" border=\"0\" alt=\"Abstract\" height=\"10pt\"> Abstract, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/pdf.gif\" border=\"0\" alt=\"Publication\" height=\"10pt\"> Publication, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/ppt.gif\" border=\"0\" alt=\"Presentation\" height=\"10pt\"> Presentation, <img decoding=\"async\" style=\"border-style: none;\" src=\"images\/bib.gif\" border=\"0\" alt=\"BibTeX Citation\" height=\"10pt\"> BibTeX Citation<\/small><\/em><\/p>\n<p><script language=\"JavaScript\">\nfunction showAbstract (text) {\n  var width  = 400;\n  var height = 400;\n  var left   = (screen.width  - width ) \/ 2;\n  var top    = (screen.height - height) \/ 2;\n  var win    = window.open('',\n                           'Abstract',\n                           'width='  + width  + ', ' + \n                           'height=' + height + ', ' +\n                           'left='   + left   + ', ' +\n                           'top='    + top    + ', ' +\n                           'toolbar=no, '     +\n                           'location=no, '    +\n                           'directories=no, ' +\n                           'status=no, '      +\n                           'menubar=no, '     +\n                           'copyhistory=no, ' +\n                           'scrollbars=yes, ' +\n                           'resizable=yes')\n  win.document.write(text);\n  win.document.close();\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Federated learning (FL) offers a collaborative framework for training foundation models (FMs) and other AI models across distributed computing infrastructures and datasets while incorporating privacy-preserving techniques to manage private-sensitive datasets. This proposal addresses the challenges inherent in adapting FL to the &#8220;pre-train&#8221; and &#8220;fine-tune&#8221; paradigms of FMs with billions or trillions of parameters. These challenges&hellip;&nbsp;<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":83,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-1188","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/1188","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1188"}],"version-history":[{"count":22,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/1188\/revisions"}],"predecessor-version":[{"id":1416,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/1188\/revisions\/1416"}],"up":[{"embeddable":true,"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=\/wp\/v2\/pages\/83"}],"wp:attachment":[{"href":"https:\/\/www.christian-engelmann.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1188"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}