Skip to search boxSkip to navigationSkip to main content

GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Host publication Subtitle

Proceedings of the Sixth European Workshop on Machine Learning and Systems

Original language

English

Pages from-to (Number of pages)

Pages 127–138 (12 pages)

Publication milestones

  • Published - 2026

Publication status

Published - 2026

Publisher

Association for Computing Machinery, United States

Book series

  • Book series name: EuroMLSys: Proceedings of the European Workshop on Machine Learning and Systems
979-8-4007-2605-7

Publication IDs

  • Scopus: 105038705741

Host publication title

EuroMLSys '26

Abstract

Collocating deep learning training tasks improves GPU utilization but risks resource contention, severe slowdowns, and out-of-memory (OOM) failures. Accurate memory estimation is essential for robust collocation, and GPU utilization estimation — a key proxy for contention — enables interference-aware scheduling.Existing GPU memory estimators span three paradigms -analytical models, CPU-side libraries, and ML-based estimators - each with distinct limitations: dependence on detailed model specifications, intrusive integration, poor generalization, and varying latency overhead. GPU heterogeneity further complicates estimation, as identical tasks can exhibit different memory footprints across hardware generations. GPU utilization remains comparatively understudied, further complicated by the non-additive utilization metrics and GPU heterogeneity.We conduct a systematic analysis of representative memory estimators from each paradigm - Horus [22], PyTorch FakeTensor [3], and our lightweight ML-based estimator - evaluating accuracy, generalizability, and overhead. We construct a synthetic dataset spanning MLPs, CNNs, and Transformers with controlled architectural variations, and train MLP- and Transformer-based estimators for memory prediction, and experiment with utilization estimation. Our evaluation reveals key tradeoffs and validates estimators against real-world unseen models. Significant challenges remain: analytical models lack generalization and cannot easily be extended to new GPU architectures or accurately reflect memory optimization savings; CPU-side libraries impose intrusive integration overhead; and both analytical and ML-based estimators rely on model specifications or computation graphs, limiting generalization across diverse model architectures and GPU hardware variants. We release all datasets, tools, and artifacts to support further research.

Funding Details

This work was funded by the Independent Research Fund Denmark’s (Danmarks Frie Forskningsfond; DFF) Sapere Aude program under grant agreement number 0171-00061B.

Related Event

Title

Computer Systems

Event type

Conference

Degree of recognition

International event

Date

27/04/2026 - 30/04/2026

Location

EdinburghUnited Kingdom