<node id="676820">
  <nid>676820</nid>
  <type>event</type>
  <uid>
    <user id="27707"><![CDATA[27707]]></user>
  </uid>
  <created>1726511000</created>
  <changed>1726511050</changed>
  <title><![CDATA[PhD Proposal by Ranjan Sarpangala Venkatesh]]></title>
  <body><![CDATA[<p><strong>Title:&nbsp;</strong>Optimizing HPC I/O Performance Over New Memory and Storage Hierarchies: A Data-Driven Approach</p><p><strong>Date</strong>: September 20th, 2024<br><strong>Time</strong>: 2:00 PM - 4:00 PM EDT<br><strong>Location</strong>: Klaus Advanced Computing Building, Conference Room 3126<br><strong>Virtual meeting</strong>: <a href="https://gatech.zoom.us/j/98262397839?pwd=fc5E5vZzIEwgH2nMHN5oa9ci1z8Q3t.1">https://gatech.zoom.us/j/98262397839?pwd=fc5E5vZzIEwgH2nMHN5oa9ci1z8Q3t.1</a></p><p><strong>Ranjan Sarpangala Venkatesh</strong><br>School of Computer Science<br>College of Computing<br>Georgia Institute of Technology</p><p><strong>Committee</strong><br>Dr. Ada Gavrilovska (advisor) - School of Computer Science, Georgia Institute of Technology<br>Dr. Greg Eisenhauer - School of Computer Science, Georgia Institute of Technology<br>Dr. Santosh Pande - School of Computer Science, Georgia Institute of Technology<br>Dr. Richard Vuduc - School of Computational Science and Engineering, Georgia Institute of Technology</p><p><strong>Abstract</strong></p><p>Multi-component HPC workflows face growing bottlenecks in data and metadata I/O due to rapid data growth. The efficiency of data movement in these workflows relies on I/O performance throughout the memory and storage hierarchy. While new memory technologies and I/O stacks offer opportunities for improvement, their distinct APIs complicate optimization, making empirical approaches necessary to balance trade-offs across components. This thesis supports data-driven methods to improve I/O performance in next-generation HPC systems.</p><p>The first part of this work evaluated various workflow configurations on systems with heterogeneous memory, showing that careful scheduling and data allocation can enhance end-to-end performance by up to 1.6x. By analyzing workflow characteristics, key elements impacting performance variability were identified, resulting in a framework for future workflow schedulers to optimize in situ workflows.</p><p>The second part focused on metadata I/O, a growing issue in large-scale workflows. Using the WarpX application and ADIOS (Adaptable I/O System) middleware, which is widely used for data management in scientific applications, it was shown that metadata I/O could account for up to 25% of total I/O time at scale. To address this, the design space of the DAOS (Distributed Asynchronous Object Storage) system was explored, specifically focusing on DAOS Key-Value and Array objects for transferring ADIOS metadata. A newly developed DAOS-based engine for ADIOS metadata I/O improved performance by 2.3x compared to the DAOS POSIX interface, effectively reducing metadata scaling bottlenecks. For the WarpX application, this reduced metadata I/O time by more than 4x, lowering overhead from 20% to just 5% of total I/O time.</p><p>As part of the proposed work, MetaBench is introduced. This suite of benchmarks will evaluate DAOS Key-Value and Array interfaces for ADIOS metadata transfer, optimized for a given HPC setup. MetaBench will analyze trade-offs, including metadata size and the number of ranks, to identify the optimal DAOS configuration. It will evaluate real applications and data patterns, providing a practical template for managing metadata transfer across HPC middleware, including ADIOS, HDF5, and PnetCDF.</p><p>These insights will contribute to the development of tools and support the HPC community by integrating DAOS engines into widely used middleware.</p>]]></body>
  <field_summary_sentence>
    <item>
      <value><![CDATA[Optimizing HPC I/O Performance Over New Memory and Storage Hierarchies: A Data-Driven Approach]]></value>
    </item>
  </field_summary_sentence>
  <field_summary>
    <item>
      <value><![CDATA[<p>Optimizing HPC I/O Performance Over New Memory and Storage Hierarchies: A Data-Driven Approach</p>]]></value>
    </item>
  </field_summary>
  <field_time>
    <item>
      <value><![CDATA[2024-09-20T14:00:00-04:00]]></value>
      <value2><![CDATA[2024-09-20T16:00:00-04:00]]></value2>
      <rrule><![CDATA[]]></rrule>
      <timezone><![CDATA[America/New_York]]></timezone>
    </item>
  </field_time>
  <field_fee>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_fee>
  <field_extras>
      </field_extras>
  <field_audience>
          <item>
        <value><![CDATA[Public]]></value>
      </item>
      </field_audience>
  <field_media>
      </field_media>
  <field_contact>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_contact>
  <field_location>
    <item>
      <value><![CDATA[Klaus Advanced Computing Building, Conference Room 3126]]></value>
    </item>
  </field_location>
  <field_sidebar>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_sidebar>
  <field_phone>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_phone>
  <field_url>
    <item>
      <url><![CDATA[]]></url>
      <title><![CDATA[]]></title>
            <attributes><![CDATA[]]></attributes>
    </item>
  </field_url>
  <field_email>
    <item>
      <email><![CDATA[]]></email>
    </item>
  </field_email>
  <field_boilerplate>
    <item>
      <nid><![CDATA[]]></nid>
    </item>
  </field_boilerplate>
  <links_related>
      </links_related>
  <files>
      </files>
  <og_groups>
          <item>221981</item>
      </og_groups>
  <og_groups_both>
          <item><![CDATA[Graduate Studies]]></item>
      </og_groups_both>
  <field_categories>
          <item>
        <tid>1788</tid>
        <value><![CDATA[Other/Miscellaneous]]></value>
      </item>
      </field_categories>
  <field_keywords>
          <item>
        <tid>102851</tid>
        <value><![CDATA[Phd proposal]]></value>
      </item>
      </field_keywords>
  <field_userdata><![CDATA[]]></field_userdata>
</node>
