Umeå University's logo

umu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Autonomous resource management for high performance datacenters
Umeå University, Faculty of Science and Technology, Department of Computing Science.
2020 (English)Doctoral thesis, comprehensive summary (Other academic)Alternative title
Autonom resurshantering för högpresterande datacenter (Swedish)
Abstract [en]

Over the last decade, new applications such as data intensive workflows have hit an inflection point in wide spread use and influenced the compute paradigm of most scientific and industrial endeavours. Data intensive workflows are highly dynamic and adaptable to resource changes, system faults, and by also allowing approximate solutions into their models. On the one hand, these dynamic characteristics require processing power and capabilities originated in cloud computing environments, and are not well supported by large High Performance Computing (HPC) infrastructures. On the other hand, cloud computing datacenters favor low latency over throughput, deeply contrasting with HPC, which enforces a centralized environment and prioritizes total computation accomplished over-time, ignoring latency entirely. Although data handling needs are predicted to increase by as much as a thousand times over the next decade, future datacenters processing power will not increase as much.

To tackle these long-term developments, this thesis proposes autonomic methods combined with novel scheduling strategies to optimize datacenter utilization while guaranteeing user defined constraints and seamlessly supporting a wide range of applications under various real operational scenarios. Leveraging upon data intensive characteristics, a library is developed to dynamically adjust the amount of resources used throughout the lifespan of a workflow, enabling elasticity for such applications in HPC datacenters. For mission critical environments where services must run even in the event of system failures, we define an adaptive controller to dynamically select the best method to perform runtime state synchronizations. We develop different hybrid extensible architectures and reinforcement learning scheduling algorithms that smoothly enable dynamic applications into HPC environments. An overall theme in this thesis is extensive experimentation in real datacenters environments. Our results show improvements in datacenter utilization and performance, achieving higher overall efficiency. Our methods also simplify operations and allow the onboarding of novel types of applications previously not supported.

Abstract [sv]

Dataintensiva workflows är en ny klass av applikationer som blivit alltmer vanliga under senaste årtiondet och har stor påverkan på hur beräkningar utförs inom flertalet forskningsområden och i industrin. Dessa dataintensiva workflows kan dynamiskt anpassa sig till ändringar i resursallokering, systemfel och kan ibland även approximera lösningar vid resursbrist. De kräver hög beräkningskraft och därtill funktionalitet som endast återfinns i datormoln och de passar därmed dåligt i dagens högpresterande datorsystem (HPC-system). Datacenter i molnet prioriterar att snabbt starta nyinkomna applikationer, vilket drastiskt skiljer sig från HPC-miljöer där hög genomströmning över tid är det främsta målet. Trots att behovet av datahantering uppskattas öka mer än tusenfallt under kommande årtioende kommer framtidens datacenter inte att ha motsvarande utveckling av beräkningskapacitet.

Denna avhandling möter dessa utmaningar genom en kombination av autonoma system och nya strategier för schedulering för att optimera utnyttjandegraden i datacenter. Detta sker utan att göra avkall på användares prestandakrav och därtill med målet att stödja ett brett spektrum av applikationer och scenarios. Ett bibliotek utvecklas för att dynamiskt anpassa resursallokering för workflows under körning, vilket innebär att även HPC-system kan stödja elastiska applikationer som tidigare bara kunde exekveras i datormoln. För miljöer med höga krav på tillgänglighet defineras en regulator för att dynamiskt anpassa hur applikationer synkroniserar tillstånd, för mer resurseffektiv aktiv replikering. Avhandlingen utvecklar även flera resurshanteringssystem baserat på schedulering med förstärkningsinlärning i syftet att förbättra stödet för dynamiska applikationer i HPC-system. Ett övergripande tema i avhandlingen är omfattande utvärderingar av de framtagna metoderna och systemen genom storskaliga experiment i verkliga datacenter. Resultaten visar förbättringar överlag av resursutnyttjande och prestanda i datacenter. De utvecklade systemen förenklar även drift och möjliggör nya typer av applikationer som tidigare ej kunnat exekveras i HPC-miljöer.

Place, publisher, year, edition, pages
Umeå: Umeå University , 2020. , p. 44
Series
Report / UMINF, ISSN 0348-0542 ; 20.03
Keywords [en]
Datacenters, high performance computing, scheduling, hybrid
National Category
Engineering and Technology
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:umu:diva-169633ISBN: 978-91-7855-286-3 (print)ISBN: 978-91-7855-287-0 (electronic)OAI: oai:DiVA.org:umu-169633DiVA, id: diva2:1423090
Public defence
2020-05-08, MIT Place Seminarierummet, MIT-byggnaden (Plan 2), Umeå, 10:00 (English)
Opponent
Supervisors
Note

New place for the public defence (wrong place in the posting sheet). 

Available from: 2020-04-17 Created: 2020-04-13 Last updated: 2020-05-25Bibliographically approved
List of papers
1. E-HPC: A Library for Elastic Resource Management in HPC Environments
Open this publication in new window or tab >>E-HPC: A Library for Elastic Resource Management in HPC Environments
Show others...
2017 (English)In: 12th Workshop on Workflows in Support of Large-Scale Science (WORKS), New York, NY, USA: Association for Computing Machinery (ACM), 2017, article id 1Conference paper, Published paper (Refereed)
Abstract [en]

Next-generation data-intensive scientific workflows need to support streaming and real-time applications with dynamic resource needs on high performance computing (HPC) platforms. The static resource allocation model on current HPC systems that was designed for monolithic MPI applications is insufficient to support the elastic resource needs of current and future workflows. In this paper, we discuss the design, implementation and evaluation of Elastic-HPC (E-HPC), an elastic framework for managing resources for scientific workflows on current HPC systems. E-HPC considers a resource slot for a workflow as an elastic window that might map to different physical resources over the duration of a workflow. Our framework uses checkpoint-restart as the underlying mechanism to migrate workflow execution across the dynamic window of resources. E-HPC provides the foundation necessary to enable dynamic resource allocation of HPC resources that are needed for streaming and real-time workflows. E-HPC has negligible overhead beyond the cost of checkpointing. Additionally, E-HPC results in decreased turnaround time of workflows compared to traditional model of resource allocation for workflows, where resources are allocated per stage of the workflow. Our evaluation shows that E-HPC improves core hour utilization for common workflow resource use patterns and provides an effective framework for elastic expansion of resources for applications with dynamic resource needs.

Place, publisher, year, edition, pages
New York, NY, USA: Association for Computing Machinery (ACM), 2017
Keywords
high performance computing, scientific workflows, resource management
National Category
Computer Sciences
Research subject
Computer Science
Identifiers
urn:nbn:se:umu:diva-142624 (URN)10.1145/3150994.3150996 (DOI)2-s2.0-85054761132 (Scopus ID)978-1-4503-5129-4 (ISBN)
Conference
The International Conference for High Performance Computing, Networking, Storage and Analysis
Available from: 2017-12-06 Created: 2017-12-06 Last updated: 2023-03-24Bibliographically approved
2. Hybrid Adaptive Checkpointing for Virtual Machine Fault Tolerance
Open this publication in new window or tab >>Hybrid Adaptive Checkpointing for Virtual Machine Fault Tolerance
Show others...
2018 (English)In: Proceedings - 2018 IEEE International Conference on Cloud Engineering, IC2E 2018 / [ed] Li J., Chandra A., Guo T., Cai Y., Institute of Electrical and Electronics Engineers (IEEE), 2018, p. 12-22Conference paper, Published paper (Refereed)
Abstract [en]

Active Virtual Machine (VM) replication is an application independent and cost-efficient mechanism for high availability and fault tolerance, with several recently proposed implementations based on checkpointing. However, these methods may suffer from large impacts on application latency, excessive resource usage overheads, and/or unpredictable behavior for varying workloads. To address these problems, we propose a hybrid approach through a Proportional-Integral (PI) controller to dynamically switch between periodic and on-demand check-pointing. Our mechanism automatically selects the method that minimizes application downtime by adapting itself to changes in workload characteristics. The implementation is based on modifications to QEMU, LibVirt, and OpenStack, to seamlessly provide fault tolerant VM provisioning and to enable the controller to dynamically select the best checkpointing mode. Our evaluation is based on experiments with a video streaming application, an e-commerce benchmark, and a software development tool. The experiments demonstrate that our adaptive hybrid approach improves both application availability and resource usage compared to static selection of a checkpointing method, with application performance gains and neglectable overheads.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2018
Keywords
Fault Tolerance, Resource Management, Checkpoint, COLO, Control Theory
National Category
Computer Systems
Research subject
Computer Science
Identifiers
urn:nbn:se:umu:diva-152033 (URN)10.1109/IC2E.2018.00023 (DOI)000759774400002 ()2-s2.0-85048315473 (Scopus ID)978-1-5386-5009-7 (ISBN)978-1-5386-5008-0 (ISBN)
Conference
2018 IEEE International Conference on Cloud Engineering (IC2E 2018), 17–20 April 2018, Orlando, Florida, USA
Available from: 2018-09-24 Created: 2018-09-24 Last updated: 2023-09-05Bibliographically approved
3. Hybrid Resource Management for HPC and Data Intensive Workloads
Open this publication in new window or tab >>Hybrid Resource Management for HPC and Data Intensive Workloads
2019 (English)In: 2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), Los Alamitos: IEEE Computer Society, 2019, p. 399-409Conference paper, Published paper (Refereed)
Abstract [en]

Traditionally, High Performance Computing (HPC) and Data Intensive (DI) workloads have been executed on separate hardware using different tools for resource and application management. With increasing convergence of these paradigms, where modern applications are composed of both types of jobs in complex workflows, this separation becomes a growing overhead and the need for a common computation platform for both application areas increases. Executing both application classes on the same hardware not only enables hybrid workflows, but can also increase the usage efficiency of the system, as often not all available hardware is fully utilized by an application. While HPC systems are typically managed in a coarse grained fashion, allocating a fixed set of resources exclusively to an application, DI systems employ a finer grained regime, enabling dynamic resource allocation and control based on application needs. On the path to full convergence, a useful and less intrusive step is a hybrid resource management system that allows the execution of DI applications on top of standard HPC scheduling systems.In this paper we present the architecture of a hybrid system enabling dual-level scheduling for DI jobs in HPC infrastructures. Our system takes advantage of real-time resource utilization monitoring to efficiently co-schedule HPC and DI applications. The architecture is easily adaptable and extensible to current and new types of distributed workloads, allowing efficient combination of hybrid workloads on HPC resources with increased job throughput and higher overall resource utilization. The architecture is implemented based on the Slurm and Mesos resource managers for HPC and DI jobs. Our experimental evaluation in a real cluster based on a set of representative HPC and DI applications demonstrate that our hybrid architecture improves resource utilization by 20%, with 12% decrease on queue makespan while still meeting all deadlines for HPC jobs.

Place, publisher, year, edition, pages
Los Alamitos: IEEE Computer Society, 2019
Keywords
Resource Management, High Performance Computing, Data Intensive Computing, Mesos, Slurm, Boostrapping
National Category
Computer Systems
Research subject
Computer Science
Identifiers
urn:nbn:se:umu:diva-155619 (URN)10.1109/CCGRID.2019.00054 (DOI)000483058700045 ()2-s2.0-85069469164 (Scopus ID)978-1-7281-0913-8 (ISBN)978-1-7281-0912-1 (ISBN)
Conference
19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGrid 2019, May 14-17, 2019, Larnaca, Cyprus
Note

Originally included in thesis in manuscript form

Available from: 2019-01-24 Created: 2019-01-24 Last updated: 2020-06-26Bibliographically approved
4. ASA - The Adaptive Scheduling Architecture
Open this publication in new window or tab >>ASA - The Adaptive Scheduling Architecture
Show others...
2020 (English)In: HPDC '20: Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing, ACM Digital Library, 2020, p. 161-165Conference paper, Published paper (Refereed)
Abstract [en]

In High Performance Computing (HPC), resources are controlled by batch systems and may not be available due to long queue waiting times, negatively impacting application deadlines. This is noticeable in low latency scientific workflows where resource planning and timely allocation are key for efficient processing. On the one hand, peak allocations guarantee the fastest possible workflows execution time, at the cost of extended queue waiting times and costly resource usage. On the other hand, dynamic allocations following specific workflow stage requirements optimizes resource usage, though it increases the total workflow makespan. To enable new scheduling strategies and features in workflows, we propose ASA: the Adaptive Scheduling Architecture, a novel scheduling method to reduce perceived queue waiting times as well as to optimize workflows resource usage. Reinforcement learning is used to estimate queue waiting times, and based on these estimates ASA pro-actively submit resource change requests, minimizing total workflow inter-stage waiting times, idle resources, and makespan. Experiments with three scientific workflows at two HPC centers show that ASA combines the best of the two aforementioned approaches, with average queue waiting time and makespan reductions of up to 10% and 2% respectively, with up to 100% prediction accuracy, while obtaining near optimal resource utilization.

Place, publisher, year, edition, pages
ACM Digital Library, 2020
Keywords
Scheduling, HPC, reinforcement learning
National Category
Engineering and Technology
Research subject
Computer Science
Identifiers
urn:nbn:se:umu:diva-169631 (URN)10.1145/3369583.3392693 (DOI)001324806800017 ()2-s2.0-85088391721 (Scopus ID)978-1-4503-7052-3 (ISBN)
Conference
HPDC '20: The 29th International Symposium on High-Performance Parallel and Distributed Computing, Stockholm, Sweden, June 23-26, 2020
Note

Originally included in thesis in manuscript form. 

Available from: 2020-04-13 Created: 2020-04-13 Last updated: 2025-04-24Bibliographically approved
5. A HPC Co-Scheduler with Reinforcement Learning
Open this publication in new window or tab >>A HPC Co-Scheduler with Reinforcement Learning
(English)Manuscript (preprint) (Other academic)
Abstract [en]

High Performance Computing (HPC) datacenters process thousands of diverse applications, supporting many scientific and business endeavours. Although users understand minimum coarse resource job requirements such as amounts of CPUs and memory, internal infrastructural utilization data and system dynamics are often visible only to cluster operators. Besides that, due to increased complexity, heuristically tweaking a batch system is even today a very challenge task. When combined with applications profiling, infrastructural data enables improvements to job scheduling, while creating space to improve Quality-of-Service (QoS) metrics such as queue waiting times and total execution times. Targeting improvements in utilization and throughput, in this paper we evaluate and propose a novel Reinforcement Learning co-scheduler algorithm that combines capacity utilization with application performance profiling. We first profile a running application by assessing its resource utilization and progress by means of a forest of decision trees, enabling our algorithm to understand the application’s resource capacity usage. We then use this information to estimate how much capacity from this ongoing allocation can be allocated for co-scheduling additional applications. Because estimations may go wrong, our algorithm has to learn and evaluate when co-scheduling decisions results in QoS degradation, such as application slowness. To overcome this, we devised a co-scheduling architecture and a handful metric to help minimizing performance degradation, enabling improvements on utilization of up to 25% even when the cluster is experiencing high demands, with 10% average queue makespan reductions when experiencing low loads.Together with the architecture, our algorithm forms the base of an application-aware co-scheduler for improved datacenter utilization and minimal performance degradation.

Keywords
Datacenters, scheduling, high performance computing, reinforcement learning
National Category
Computer Systems
Research subject
Computer Science
Identifiers
urn:nbn:se:umu:diva-169632 (URN)
Available from: 2020-04-13 Created: 2020-04-13 Last updated: 2020-04-14

Open Access in DiVA

fulltext(1222 kB)732 downloads
File information
File name FULLTEXT01.pdfFile size 1222 kBChecksum SHA-512
424204091206f337b2fa260cd2c73b628edbd183edf97b950aa4b990884487e3fd6ba679c239041bc425d003fd26449e493fb0b0f1316b45e3748fbde5b34c98
Type fulltextMimetype application/pdf
spikblad(61 kB)68 downloads
File information
File name SPIKBLAD03.pdfFile size 61 kBChecksum SHA-512
50d85b4dc17846ee804dc694e34e1e2830e2b3ad0651536186fa82ca66def2d72b912e489c5c105bdb5bd3f2bff65e1a1fc297fe76fbfbf1c33c3484d3b6d572
Type spikbladMimetype application/pdf
errata(1181 kB)85 downloads
File information
File name ERRATA04.pdfFile size 1181 kBChecksum SHA-512
b7d0fa740b8fe6d50d7716357ee49086c879d97cf9874547a327e33acbb233b2d3df942f1970bfcdc35b6cb85a907f55aea8852f1804115ab93ab2fcc95ad020
Type errataMimetype application/pdf

Authority records

Souza, Abel

Search in DiVA

By author/editor
Souza, Abel
By organisation
Department of Computing Science
Engineering and Technology

Search outside of DiVA

GoogleGoogle Scholar
Total: 733 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 2247 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf