No GPUs Were Harmed: Predicting DL Training Memory Footprints Before Launch

  • 2 October 2026
  • 2pm - 2:50pm
  • Dr Yehia Elkhatib

Abstract

Training and fine-tuning large AI models requires immense GPU memory, but cloud providers and researchers face a frustrating dilemma: underestimating a job's memory footprint causes GPU crashes due to Out-of-Memory (OOM) errors; overprovisioning lets thousands of dollars' worth of GPU hardware sit half-idle. Because modern software allocators manage memory dynamically behind the scenes, predicting whether a model will fit before launching it has traditionally been almost impossible without expensive trial-and-error profiling.

This talk walks through methods to accurately forecast peak memory demands without risking hardware crashes or code modifications. We explore 2 systems: one that runs a lightweight 'dry run' on standard, low-cost CPU hardware and simulates how GPU memory allocators behave, and another that uses a symbolic representation to understand how model families reuse memory layers without running any code (even on a CPU). This talk plots the path from trial-and-error, replay, and symbolic maths; effectively demonstrating how moving from hardware profiling to symbolic prediction enables zero-execution, crash-free AI computing.

Speaker

Dr Yehia Elkhatib is a Reader in Computer Science at the University of Glasgow. His research focuses on developing data-driven tools to optimise complex distributed systems, with particular emphasis on addressing resource utilisation challenges from a user-centric perspective. Through his research, he addresses decision-making challenges across cloud, edge, IoT, and robotic systems. He is the TPC co-chair of IEEE Cloud 2025 and 2026, and an associate editor of IEEE Transactions on Services Computing (TSC).

Contact and booking details

Booking required?
No