HiPC 2026 Tutorial

ARISE

AI Research through Infrastructure Simulation and Evaluation

A hands-on tutorial on simulation-driven methodologies for designing and evaluating next-generation AI infrastructure โ€” communication fabrics, memory systems, and large-scale distributed training & inference.

Format Hands-on Tutorial
Focus AI Infra Simulation at Scale

Abstract

Why this tutorial matters

The rapid growth of foundation models and large-scale AI systems has transformed modern computing infrastructure. As model sizes and cluster scales continue to grow, system performance is increasingly constrained not by compute alone, but by communication, memory capacity, memory bandwidth, collective operations, and interconnect fabrics. Emerging technologies such as Ultra Ethernet, InfiniBand, NVLink, UALink, CXL-based memory expansion, and memory-semantic fabrics are redefining the architecture of AI clusters and introducing new research challenges spanning systems, networking, and computer architecture.

This tutorial introduces simulation-driven methodologies for studying and designing next-generation AI infrastructure. Participants will learn the fundamentals of distributed AI training and inference, understand the role of communication and memory bottlenecks, and gain hands-on experience using Astra-Sim and related simulation frameworks. The tutorial further covers emerging research directions in AI fabrics, memory disaggregation, large-scale model serving, and memory-semantic networking.

The tutorial combines conceptual foundations, practical demonstrations, and guided hands-on exercises, enabling researchers and practitioners to evaluate AI infrastructure designs without requiring access to expensive production-scale clusters.

Objectives

  • Introduce the architecture of modern AI infrastructure.
  • Explain communication and memory bottlenecks in large-scale AI systems.
  • Provide hands-on experience using Astra-Sim and LLM-Serving-Sim for modeling AI workloads.
  • Discuss emerging research opportunities in memory-semantic fabrics, AI networking, and disaggregated AI infrastructure, and demonstrate how simulation can be used for such research problems.

Intended Audience

  • Graduate students and researchers in HPC, systems, networking, and computer architecture.
  • Industry practitioners working on AI infrastructure and large-scale training and inference.

Prerequisites & Level

  • Basic computer architecture
  • Basic networking concepts
  • Familiarity with Linux
  • No prior experience with AI infrastructure simulators is required.

Tutorial Topics

Four parts, one design space

Part I

AI Research Using Simulation Tools

  • Why simulation-driven research matters for AI infrastructure
  • Overview of Astra-Sim and related simulation frameworks
  • Motivating case studies from industry-scale AI clusters
Part II

Foundations Recap: Collectives & Parallelization

  • Data, Tensor, Pipeline, and Expert Parallelism
  • Collective communication primitives
  • AI communication fabrics: InfiniBand, Ultra Ethernet, NVLink / UALink, RDMA
Part III

Hands-on: Fabric Performance Evaluation, UET vs RoCEv2

  • Model scale-up/scale-out AI cluster topologies under UET and RoCEv2 transports
  • Measure and compare collective communication latency, throughput, and congestion behavior
  • Discuss trade-offs for next-generation AI fabric design choices
Part IV

Research Directions in AI Infrastructure

  • Emerging research opportunities in memory-semantic fabrics and AI networking
  • Distributed AI systems and industry perspectives
  • Open problems in disaggregated AI infrastructure

Tutorial Schedule

Agenda

Exact session timings are TBD and subject to final confirmation. Additional presenters may be included depending on the final tutorial team composition.

Time Topic Presenter(s)
Session 1
Timing TBD
AI Research Using Simulation Tools Abed Mohammad Kamaluddin (Marvell)
Session 2
Timing TBD
Quick Recap: Collectives & Parallelization Strategies Ramanjeet Singh (Marvell), Sriram Vatala (Marvell)
Session 3
Timing TBD
Hands-on: Fabric Performance Evaluation โ€” UET vs RoCEv2
  • Model scale-up/scale-out AI cluster topologies under Ultra Ethernet (UET) and RoCEv2 transports
  • Measure and compare collective communication latency, throughput, and congestion behavior
  • Discuss trade-offs for next-generation AI fabric design choices
Ramanjeet Singh (Marvell), Sriram Vatala (Marvell), Hemant (Marvell)
Session 4
Timing TBD
Research Directions in AI Infrastructure Dr Rinku Shah

Tutorial Presenters

Meet the Team

Rinku Shah

Assistant Professor, IIIT-Delhi

Rinku is currently an Assistant Professor in the CSE department at Indraprastha Institute of Information Technology, Delhi (IIITD). She earned her PhD from IIT Bombay in February 2021. Her research interests span computer networks, distributed systems, AI infrastructure, programmable networking, cloud computing, and security. She is particularly interested in designing abstractions and architectures for flexible, high-performance, and secure networking solutions that meet the needs of next-generation distributed systems. She has published her research at reputed conference venues such as ACM CoNEXT, IEEE ICNP, IEEE/IFIP DSN, ACM SOSR, and IEEE/IFIP Networking.

Abed Mohammad Kamaluddin

Director, Marvell Technology

Abed Mohammad Kamaluddin is a Director at Marvell Technology, where he drives solution and platform initiatives for next-generation AI infrastructure, networking, and memory-semantic fabrics. He has extensive experience in large-scale systems, networking, AI fabrics, distributed AI infrastructure, and semiconductor platforms. He has played a leading role in industry-academia collaborations focused on AI systems and networking, including co-developing and co-teaching advanced courses on Networks for AI/ML Systems. He is actively involved in community-building efforts around AI infrastructure and serves on the organizing teams of emerging workshops focused on AI fabrics, memory-semantic networking, and large-scale AI systems research.

Sriram Vatala

Senior Staff Engineer, Marvell Technology

Sriram Vatala is a Senior Staff Engineer at Marvell Technology with 9 years of industry experience in user-space packet processing, dataplane, and IPSec systems, with a focus on high-performance networking software for AI infrastructure.

Ramanjeet Singh

Software Engineer, Marvell Technology

Ramanjeet Singh is a Software Engineer at Marvell Technology with about 2 years of industry experience. He holds an undergraduate degree from IIIT-Delhi and works on AI infrastructure and networking systems.

Hemant

Senior Engineer, Marvell Technology

Hemant is a Senior Engineer at Marvell Technology with about 2 years of industry experience. He holds a postgraduate degree from BITS Pilani and works on AI infrastructure and networking systems.

Additional presenters may be included depending on the final tutorial team composition.

Tutorial Website

Materials & Resources

Upon acceptance, this site will host the full set of tutorial materials below. All materials will remain publicly available after the conference to maximize community impact.

๐Ÿ“ฝ๏ธ

Slides

Presentation decks for all three parts of the tutorial.

Coming soon
๐Ÿงช

Hands-on Lab Guides

Step-by-step guides for the Astra-Sim hands-on sessions.

Coming soon
โš™๏ธ

Software Setup

Installation and environment setup instructions for the simulators.

Coming soon
๐Ÿ“š

References & Datasets

Reference readings and datasets used throughout the tutorial.

Coming soon