Umeå University's logo

umu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
A HPC Co-scheduler with Reinforcement Learning
University of Massachusetts Amherst, Amherst, USA.ORCID iD: 0000-0001-6952-1195
Uppsala University, Uppsala, Sweden.
Umeå University, Faculty of Science and Technology, Department of Computing Science.
2021 (English)In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Springer, 2021, p. 126-148Conference paper, Published paper (Refereed)
Abstract [en]

Although High Performance Computing (HPC) users understand basic resource requirements such as the number of CPUs and memory limits, internal infrastructural utilization data is exclusively leveraged by cluster operators, who use it to configure batch schedulers. This task is challenging and increasingly complex due to ever larger cluster scales and heterogeneity of modern scientific workflows. As a result, HPC systems achieve low utilization with long job completion times (makespans). To tackle these challenges, we propose a co-scheduling algorithm based on an adaptive reinforcement learning algorithm, where application profiling is combined with cluster monitoring. The resulting cluster scheduler matches resource utilization to application performance in a fine-grained manner (i.e., operating system level). As opposed to nominal allocations, we apply decision trees to model applications’ actual resource usage, which are used to estimate how much resource capacity from one allocation can be co-allocated to additional applications. Our algorithm learns from incorrect co-scheduling decisions and adapts from changing environment conditions, and evaluates when such changes cause resource contention that impacts quality of service metrics such as jobs slowdowns. We integrate our algorithm in an HPC resource manager that combines Slurm and Mesos for job scheduling and co-allocation, respectively. Our experimental evaluation performed in a dedicated cluster executing a mix of four real different scientific workflows demonstrates improvements on cluster utilization of up to 51% even in high load scenarios, with 55% average queue makespan reductions under low loads.

Place, publisher, year, edition, pages
Springer, 2021. p. 126-148
Series
Lecture Notes in Computer Science (LNCS), ISSN 0302-9743, E-ISSN 1611-3349 ; 12985
Keywords [en]
Adaptive reinforcement learning, Co-scheduling, Datacenters, High performance computing
National Category
Computer Systems
Identifiers
URN: urn:nbn:se:umu:diva-188961DOI: 10.1007/978-3-030-88224-2_7ISI: 000869960400007Scopus ID: 2-s2.0-85117482466ISBN: 978-3-030-88223-5 (print)ISBN: 978-3-030-88224-2 (electronic)OAI: oai:DiVA.org:umu-188961DiVA, id: diva2:1606848
Conference
JSSPP 2021, 24th International Workshop on Job Scheduling Strategies for Parallel Processing, Virtual, May 21, 2021
Note

Also part of the Theoretical Computer Science and General Issues book sub series (LNTCS, volume 12985)

Available from: 2021-10-28 Created: 2021-10-28 Last updated: 2023-09-05Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Souza, AbelTordsson, Johan

Search in DiVA

By author/editor
Souza, AbelTordsson, Johan
By organisation
Department of Computing Science
Computer Systems

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 356 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf