An Analysis of Collocation on GPUs for Deep Learning Training
Publikation:
Konference artikel i Proceeding eller bog/rapport kapitel
Konferencebidrag i proceedings
Peer-reviewOpen Access
Publikation information
Produktionstype
Publikation:
Konference artikel i Proceeding eller bog/rapport kapitel
Konferencebidrag i proceedings
Peer-reviewOriginalsprog
EngelskSider fra-til (Antal sider)
Sider 81-90 (10 sider)Publikationsmilepæle
- Udgivet - 22/04/2024
Publikationsstatus
Udgivet - 22/04/2024
Forlag
Association for Computing Machinery, USAISBN (Trykt)
9798400705410Publication IDs
- Scopus: 85192265987
Titel på værtspublikation
Proceedings of the 4th Workshop on Machine Learning and Systems, EuroMLSys 2024, Athens, Greece, 22 April 2024Resume
Deep learning training is an expensive process that extensively uses GPUs. However, not all model training saturates modern powerful GPUs. To create guidelines for such cases,
this paper examines the performance of the different collocation methods available on NVIDIA GPUs: naïvely submitting multiple processes on the same GPU using multiple streams,
utilizing Multi-Process Service (MPS), and enabling the MultiInstance GPU (MIG). Our results demonstrate that collocating multiple model training runs yields significant benefits, leading to up to three times training throughput despite increased epoch time. On the other hand, the aggregate memory footprint and compute needs of the models trained in parallel must fit the available memory and compute resources of the GPU. MIG can be beneficial thanks to its interference-free partitioning but can suffer from sub-optimal GPU utilization with dynamic or mixed workloads. In general, we recommend MPS as the best-performing and most flexible form of collocation for a single user submitting training jobs.
this paper examines the performance of the different collocation methods available on NVIDIA GPUs: naïvely submitting multiple processes on the same GPU using multiple streams,
utilizing Multi-Process Service (MPS), and enabling the MultiInstance GPU (MIG). Our results demonstrate that collocating multiple model training runs yields significant benefits, leading to up to three times training throughput despite increased epoch time. On the other hand, the aggregate memory footprint and compute needs of the models trained in parallel must fit the available memory and compute resources of the GPU. MIG can be beneficial thanks to its interference-free partitioning but can suffer from sub-optimal GPU utilization with dynamic or mixed workloads. In general, we recommend MPS as the best-performing and most flexible form of collocation for a single user submitting training jobs.
Metrikker
PlumX, åbner i en ny fane
Hentninger
7
Citationer
10
Adgang til dokumenter
Relateret event
Titel
Workshop on Machine Learning and Systems
Begivenhedstype
WorkshopGrad af anerkendelse
International begivenhedDato
22/04/2024 - 22/04/2024Lokation
AthensGrækenland
