Umeå University's logo

umu.sePublikasjoner
Endre søk
Link to record
Permanent link

Direct link
Publikasjoner (10 av 12) Visa alla publikasjoner
Wen, Y., Townend, P., Östberg, P.-O., Souza, A. & Courageux-Sudan, C. (2025). A decentralized microservice scheduling approach using service mesh in cloud-edge systems. In: Lisa O’Conner (Ed.), 2025 IEEE international conference on joint cloud computing: proceedings. Paper presented at IEEE JCC 2025 – The 16th IEEE International Conference on JointCloud Computing (part of IEEE CISOSE 2025), Tuscon, Arizona, USA, July 21-24, 2025 (pp. 52-60). IEEE Computer Society
Åpne denne publikasjonen i ny fane eller vindu >>A decentralized microservice scheduling approach using service mesh in cloud-edge systems
Vise andre…
2025 (engelsk)Inngår i: 2025 IEEE international conference on joint cloud computing: proceedings / [ed] Lisa O’Conner, IEEE Computer Society, 2025, s. 52-60Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

As microservice-based systems scale across thecloud-edge continuum, traditional centralized scheduling mecha-nisms increasingly struggle with latency, coordination overhead,and fault tolerance. This paper presents a new architectural di-rection: leveraging service mesh sidecar proxies as decentralized,in-situ schedulers to enable scalable, low-latency coordination inlarge-scale, cloud-native environments. We propose embeddinglightweight, autonomous scheduling logic into each sidecar, allow-ing scheduling decisions to be made locally without centralizedcontrol. This approach leverages the growing maturity of servicemesh infrastructures, which support programmable distributedtraffic management. We describe the design of such an archi-tecture and present initial results demonstrating its scalabilitypotential in terms of response time and latency under varyingrequest rates. Rather than delivering a finalized scheduling algo-rithm, this paper presents a system-level architectural directionand preliminary evidence to support its scalability potential.

sted, utgiver, år, opplag, sider
IEEE Computer Society, 2025
Emneord
Microservice-based systems, Service mesh, Decentralized scheduling, Sidecar proxy, Scalability, Latency, Distributed systems
HSV kategori
Forskningsprogram
datalogi; datorteknik
Identifikatorer
urn:nbn:se:umu:diva-241674 (URN)10.1109/JCC67032.2025.00012 (DOI)2-s2.0-105016245878 (Scopus ID)979-8-3315-8915-8 (ISBN)
Konferanse
IEEE JCC 2025 – The 16th IEEE International Conference on JointCloud Computing (part of IEEE CISOSE 2025), Tuscon, Arizona, USA, July 21-24, 2025
Tilgjengelig fra: 2025-06-28 Laget: 2025-06-28 Sist oppdatert: 2025-10-14bibliografisk kontrollert
Villarroel, B., Pelckmans, K., Solano, E., Laaksoharju, M., Souza, A., Dom, O. N., . . . Ward, M. J. (2022). Launching the VASCO Citizen Science Project. Universe, 8(11), Article ID 561.
Åpne denne publikasjonen i ny fane eller vindu >>Launching the VASCO Citizen Science Project
Vise andre…
2022 (engelsk)Inngår i: Universe, E-ISSN 2218-1997, Vol. 8, nr 11, artikkel-id 561Artikkel i tidsskrift (Fagfellevurdert) Published
Abstract [en]

The Vanishing & Appearing Sources during a Century of Observations (VASCO) project investigates astronomical surveys spanning a time interval of 70 years, searching for unusual and exotic transients. We present herein the VASCO Citizen Science Project, which can identify unusual candidates driven by three different approaches: hypothesis, exploratory, and machine learning, which is particularly useful for SETI searches. To address the big data challenge, VASCO combines three methods: the Virtual Observatory, user-aided machine learning, and visual inspection through citizen science. Here we demonstrate the citizen science project and its improved candidate selection process, and we give a progress report. We also present the VASCO citizen science network led by amateur astronomy associations mainly located in Algeria, Cameroon, and Nigeria. At the moment of writing, the citizen science project has carefully examined 15,593 candidate image pairs in the data (ca. 10% of the candidates), and has so far identified 798 objects classified as "vanished". The most interesting candidates will be followed up with optical and infrared imaging, together with the observations by the most potent radio telescopes.

sted, utgiver, år, opplag, sider
MDPI, 2022
Emneord
citizen science, SETI, surveys, transients
HSV kategori
Identifikatorer
urn:nbn:se:umu:diva-201224 (URN)10.3390/universe8110561 (DOI)000881392200001 ()2-s2.0-85141787535 (Scopus ID)
Forskningsfinansiär
Swedish Research Council, 2017-06372NordForsk
Tilgjengelig fra: 2022-12-12 Laget: 2022-12-12 Sist oppdatert: 2022-12-12bibliografisk kontrollert
Souza, A., Pelckmans, K. & Tordsson, J. (2021). A HPC Co-scheduler with Reinforcement Learning. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics): . Paper presented at JSSPP 2021, 24th International Workshop on Job Scheduling Strategies for Parallel Processing, Virtual, May 21, 2021 (pp. 126-148). Springer
Åpne denne publikasjonen i ny fane eller vindu >>A HPC Co-scheduler with Reinforcement Learning
2021 (engelsk)Inngår i: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Springer, 2021, s. 126-148Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Although High Performance Computing (HPC) users understand basic resource requirements such as the number of CPUs and memory limits, internal infrastructural utilization data is exclusively leveraged by cluster operators, who use it to configure batch schedulers. This task is challenging and increasingly complex due to ever larger cluster scales and heterogeneity of modern scientific workflows. As a result, HPC systems achieve low utilization with long job completion times (makespans). To tackle these challenges, we propose a co-scheduling algorithm based on an adaptive reinforcement learning algorithm, where application profiling is combined with cluster monitoring. The resulting cluster scheduler matches resource utilization to application performance in a fine-grained manner (i.e., operating system level). As opposed to nominal allocations, we apply decision trees to model applications’ actual resource usage, which are used to estimate how much resource capacity from one allocation can be co-allocated to additional applications. Our algorithm learns from incorrect co-scheduling decisions and adapts from changing environment conditions, and evaluates when such changes cause resource contention that impacts quality of service metrics such as jobs slowdowns. We integrate our algorithm in an HPC resource manager that combines Slurm and Mesos for job scheduling and co-allocation, respectively. Our experimental evaluation performed in a dedicated cluster executing a mix of four real different scientific workflows demonstrates improvements on cluster utilization of up to 51% even in high load scenarios, with 55% average queue makespan reductions under low loads.

sted, utgiver, år, opplag, sider
Springer, 2021
Serie
Lecture Notes in Computer Science (LNCS), ISSN 0302-9743, E-ISSN 1611-3349 ; 12985
Emneord
Adaptive reinforcement learning, Co-scheduling, Datacenters, High performance computing
HSV kategori
Identifikatorer
urn:nbn:se:umu:diva-188961 (URN)10.1007/978-3-030-88224-2_7 (DOI)000869960400007 ()2-s2.0-85117482466 (Scopus ID)978-3-030-88223-5 (ISBN)978-3-030-88224-2 (ISBN)
Konferanse
JSSPP 2021, 24th International Workshop on Job Scheduling Strategies for Parallel Processing, Virtual, May 21, 2021
Merknad

Also part of the Theoretical Computer Science and General Issues book sub series (LNTCS, volume 12985)

Tilgjengelig fra: 2021-10-28 Laget: 2021-10-28 Sist oppdatert: 2023-09-05bibliografisk kontrollert
Gutierrez, F., Beedkar, K., Souza, A. & Markl, V. (2021). AdCom: Adaptive combiner for streaming aggregations. In: Velegrakis Y.; Velegrakis Y.; Zeinalipour D.; Chrysanthis P.K.; Chrysanthis P.K.; Guerra F. (Ed.), Advances in Database Technology - EDBT: . Paper presented at Advances in Database Technology - 24th International Conference on Extending Database Technology, EDBT 2021, Virtual, March 23-26, 2021. (pp. 403-414). OpenProceedings, 2021-March
Åpne denne publikasjonen i ny fane eller vindu >>AdCom: Adaptive combiner for streaming aggregations
2021 (engelsk)Inngår i: Advances in Database Technology - EDBT / [ed] Velegrakis Y.; Velegrakis Y.; Zeinalipour D.; Chrysanthis P.K.; Chrysanthis P.K.; Guerra F., OpenProceedings, 2021, Vol. 2021-March, s. 403-414Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Continuous applications such as device monitoring and anomaly detection often require real-time aggregated statistics over unbounded data streams. While existing stream processing systems such as Flink, Spark, and Storm support processing of streaming aggregations, their optimizations are limited with respect to the dynamic nature of the data, and therefore are suboptimal when the workload changes and/or when there is data skew. In this paper we present AdCom, which is an adaptive combiner for stream processing engines. The use of AdCom in aggregation queries enables pre-aggregating tuples upstream (i.e., before data shuffling) followed by global aggregation downstream. In contrast to existing approaches, AdCom can automatically adjust the number of tuples to pre-aggregate depending on the data rate and available network. Our experimental study using real-world streaming workloads shows that using AdCom leads to 2.5-9× higher sustainable throughput without compromising latency.

sted, utgiver, år, opplag, sider
OpenProceedings, 2021
Serie
Advances in Database Technology, E-ISSN 2367-2005
HSV kategori
Forskningsprogram
data- och systemvetenskap
Identifikatorer
urn:nbn:se:umu:diva-187179 (URN)10.5441/002/edbt.2021.43 (DOI)2-s2.0-85113710327 (Scopus ID)9783893180844 (ISBN)
Konferanse
Advances in Database Technology - 24th International Conference on Extending Database Technology, EDBT 2021, Virtual, March 23-26, 2021.
Forskningsfinansiär
EU, Horizon 2020, 765452
Tilgjengelig fra: 2021-09-06 Laget: 2021-09-06 Sist oppdatert: 2021-09-06bibliografisk kontrollert
Souza, A., Pelckmans, K., Ghoshal, D., Ramakrishnan, L. & Tordsson, J. (2020). ASA - The Adaptive Scheduling Architecture. In: HPDC '20: Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing: . Paper presented at HPDC '20: The 29th International Symposium on High-Performance Parallel and Distributed Computing, Stockholm, Sweden, June 23-26, 2020 (pp. 161-165). ACM Digital Library
Åpne denne publikasjonen i ny fane eller vindu >>ASA - The Adaptive Scheduling Architecture
Vise andre…
2020 (engelsk)Inngår i: HPDC '20: Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing, ACM Digital Library, 2020, s. 161-165Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

In High Performance Computing (HPC), resources are controlled by batch systems and may not be available due to long queue waiting times, negatively impacting application deadlines. This is noticeable in low latency scientific workflows where resource planning and timely allocation are key for efficient processing. On the one hand, peak allocations guarantee the fastest possible workflows execution time, at the cost of extended queue waiting times and costly resource usage. On the other hand, dynamic allocations following specific workflow stage requirements optimizes resource usage, though it increases the total workflow makespan. To enable new scheduling strategies and features in workflows, we propose ASA: the Adaptive Scheduling Architecture, a novel scheduling method to reduce perceived queue waiting times as well as to optimize workflows resource usage. Reinforcement learning is used to estimate queue waiting times, and based on these estimates ASA pro-actively submit resource change requests, minimizing total workflow inter-stage waiting times, idle resources, and makespan. Experiments with three scientific workflows at two HPC centers show that ASA combines the best of the two aforementioned approaches, with average queue waiting time and makespan reductions of up to 10% and 2% respectively, with up to 100% prediction accuracy, while obtaining near optimal resource utilization.

sted, utgiver, år, opplag, sider
ACM Digital Library, 2020
Emneord
Scheduling, HPC, reinforcement learning
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
urn:nbn:se:umu:diva-169631 (URN)10.1145/3369583.3392693 (DOI)001324806800017 ()2-s2.0-85088391721 (Scopus ID)978-1-4503-7052-3 (ISBN)
Konferanse
HPDC '20: The 29th International Symposium on High-Performance Parallel and Distributed Computing, Stockholm, Sweden, June 23-26, 2020
Merknad

Originally included in thesis in manuscript form. 

Tilgjengelig fra: 2020-04-13 Laget: 2020-04-13 Sist oppdatert: 2025-04-24bibliografisk kontrollert
Souza, A. (2020). Autonomous resource management for high performance datacenters. (Doctoral dissertation). Umeå: Umeå University
Åpne denne publikasjonen i ny fane eller vindu >>Autonomous resource management for high performance datacenters
2020 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Alternativ tittel[sv]
Autonom resurshantering för högpresterande datacenter
Abstract [en]

Over the last decade, new applications such as data intensive workflows have hit an inflection point in wide spread use and influenced the compute paradigm of most scientific and industrial endeavours. Data intensive workflows are highly dynamic and adaptable to resource changes, system faults, and by also allowing approximate solutions into their models. On the one hand, these dynamic characteristics require processing power and capabilities originated in cloud computing environments, and are not well supported by large High Performance Computing (HPC) infrastructures. On the other hand, cloud computing datacenters favor low latency over throughput, deeply contrasting with HPC, which enforces a centralized environment and prioritizes total computation accomplished over-time, ignoring latency entirely. Although data handling needs are predicted to increase by as much as a thousand times over the next decade, future datacenters processing power will not increase as much.

To tackle these long-term developments, this thesis proposes autonomic methods combined with novel scheduling strategies to optimize datacenter utilization while guaranteeing user defined constraints and seamlessly supporting a wide range of applications under various real operational scenarios. Leveraging upon data intensive characteristics, a library is developed to dynamically adjust the amount of resources used throughout the lifespan of a workflow, enabling elasticity for such applications in HPC datacenters. For mission critical environments where services must run even in the event of system failures, we define an adaptive controller to dynamically select the best method to perform runtime state synchronizations. We develop different hybrid extensible architectures and reinforcement learning scheduling algorithms that smoothly enable dynamic applications into HPC environments. An overall theme in this thesis is extensive experimentation in real datacenters environments. Our results show improvements in datacenter utilization and performance, achieving higher overall efficiency. Our methods also simplify operations and allow the onboarding of novel types of applications previously not supported.

Abstract [sv]

Dataintensiva workflows är en ny klass av applikationer som blivit alltmer vanliga under senaste årtiondet och har stor påverkan på hur beräkningar utförs inom flertalet forskningsområden och i industrin. Dessa dataintensiva workflows kan dynamiskt anpassa sig till ändringar i resursallokering, systemfel och kan ibland även approximera lösningar vid resursbrist. De kräver hög beräkningskraft och därtill funktionalitet som endast återfinns i datormoln och de passar därmed dåligt i dagens högpresterande datorsystem (HPC-system). Datacenter i molnet prioriterar att snabbt starta nyinkomna applikationer, vilket drastiskt skiljer sig från HPC-miljöer där hög genomströmning över tid är det främsta målet. Trots att behovet av datahantering uppskattas öka mer än tusenfallt under kommande årtioende kommer framtidens datacenter inte att ha motsvarande utveckling av beräkningskapacitet.

Denna avhandling möter dessa utmaningar genom en kombination av autonoma system och nya strategier för schedulering för att optimera utnyttjandegraden i datacenter. Detta sker utan att göra avkall på användares prestandakrav och därtill med målet att stödja ett brett spektrum av applikationer och scenarios. Ett bibliotek utvecklas för att dynamiskt anpassa resursallokering för workflows under körning, vilket innebär att även HPC-system kan stödja elastiska applikationer som tidigare bara kunde exekveras i datormoln. För miljöer med höga krav på tillgänglighet defineras en regulator för att dynamiskt anpassa hur applikationer synkroniserar tillstånd, för mer resurseffektiv aktiv replikering. Avhandlingen utvecklar även flera resurshanteringssystem baserat på schedulering med förstärkningsinlärning i syftet att förbättra stödet för dynamiska applikationer i HPC-system. Ett övergripande tema i avhandlingen är omfattande utvärderingar av de framtagna metoderna och systemen genom storskaliga experiment i verkliga datacenter. Resultaten visar förbättringar överlag av resursutnyttjande och prestanda i datacenter. De utvecklade systemen förenklar även drift och möjliggör nya typer av applikationer som tidigare ej kunnat exekveras i HPC-miljöer.

sted, utgiver, år, opplag, sider
Umeå: Umeå University, 2020. s. 44
Serie
Report / UMINF, ISSN 0348-0542 ; 20.03
Emneord
Datacenters, high performance computing, scheduling, hybrid
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
urn:nbn:se:umu:diva-169633 (URN)978-91-7855-286-3 (ISBN)978-91-7855-287-0 (ISBN)
Disputas
2020-05-08, MIT Place Seminarierummet, MIT-byggnaden (Plan 2), Umeå, 10:00 (engelsk)
Opponent
Veileder
Merknad

New place for the public defence (wrong place in the posting sheet). 

Tilgjengelig fra: 2020-04-17 Laget: 2020-04-13 Sist oppdatert: 2020-05-25bibliografisk kontrollert
Villarroel, B., Pelckmans, K., Solano, E., Laaksoharju, M., Souza, A., Dom, O. N., . . . Ward, M. J. (2020). The VASCO project: 100 red transients and their follow up. In: Proceedings of the International Astronautical Congress, IAC: 71st International Astronautical Congress, IAC 2020. Paper presented at 71st International Astronautical Congress, IAC 2020; Virtual, October 12-14, 2020. International Astronautical Federation, IAF, Article ID 166680.
Åpne denne publikasjonen i ny fane eller vindu >>The VASCO project: 100 red transients and their follow up
Vise andre…
2020 (engelsk)Inngår i: Proceedings of the International Astronautical Congress, IAC: 71st International Astronautical Congress, IAC 2020, International Astronautical Federation, IAF , 2020, artikkel-id 166680Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

The Vanishing & Appearing Sources during a Century of Observations (VASCO) project investigates astronomical surveys spanning a 70 years time interval, searching for unusual and exotic transients. We present herein the VASCO Citizen Science Project, that uses three different approaches to the identification of unusual transients in a given set of candidates: hypothesis-driven, exploratory-driven and machine learning-driven (which is of particular benefit for SETI searches). To address the big data challenge, VASCO combines methods from the Virtual Observatory, a user-aided machine learning and visual inspection through citizen science. In this article, we demonstrate the citizen science project, the new and improved candidate selection process and give a progress report. We also present the VASCO citizen science network led by amateur astronomy associations mainly located in Algeria, Cameroon and Nigeria. At the moment of writing, the citizen science project has carefully examined 12,000 candidate image pairs in the data, and has so far identified 713 objects classified as “vanished”. The most interesting candidates will be followed up with optical and infrared imaging, together with the observations by the most potent radio telescopes.

sted, utgiver, år, opplag, sider
International Astronautical Federation, IAF, 2020
Serie
Proceedings of the International Astronautical Congress, IAC, ISSN 0074-1795
Emneord
Catalogs, Extraterrestrial intelligence, Miscellaneous, Surveys, Transient
HSV kategori
Identifikatorer
urn:nbn:se:umu:diva-181050 (URN)2-s2.0-85100949904 (Scopus ID)
Konferanse
71st International Astronautical Congress, IAC 2020; Virtual, October 12-14, 2020
Tilgjengelig fra: 2021-03-05 Laget: 2021-03-05 Sist oppdatert: 2021-04-16bibliografisk kontrollert
Souza, A., Rezaei, M., Laure, E. & Tordsson, J. (2019). Hybrid Resource Management for HPC and Data Intensive Workloads. In: 2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID): . Paper presented at 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGrid 2019, May 14-17, 2019, Larnaca, Cyprus (pp. 399-409). Los Alamitos: IEEE Computer Society
Åpne denne publikasjonen i ny fane eller vindu >>Hybrid Resource Management for HPC and Data Intensive Workloads
2019 (engelsk)Inngår i: 2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), Los Alamitos: IEEE Computer Society, 2019, s. 399-409Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Traditionally, High Performance Computing (HPC) and Data Intensive (DI) workloads have been executed on separate hardware using different tools for resource and application management. With increasing convergence of these paradigms, where modern applications are composed of both types of jobs in complex workflows, this separation becomes a growing overhead and the need for a common computation platform for both application areas increases. Executing both application classes on the same hardware not only enables hybrid workflows, but can also increase the usage efficiency of the system, as often not all available hardware is fully utilized by an application. While HPC systems are typically managed in a coarse grained fashion, allocating a fixed set of resources exclusively to an application, DI systems employ a finer grained regime, enabling dynamic resource allocation and control based on application needs. On the path to full convergence, a useful and less intrusive step is a hybrid resource management system that allows the execution of DI applications on top of standard HPC scheduling systems.In this paper we present the architecture of a hybrid system enabling dual-level scheduling for DI jobs in HPC infrastructures. Our system takes advantage of real-time resource utilization monitoring to efficiently co-schedule HPC and DI applications. The architecture is easily adaptable and extensible to current and new types of distributed workloads, allowing efficient combination of hybrid workloads on HPC resources with increased job throughput and higher overall resource utilization. The architecture is implemented based on the Slurm and Mesos resource managers for HPC and DI jobs. Our experimental evaluation in a real cluster based on a set of representative HPC and DI applications demonstrate that our hybrid architecture improves resource utilization by 20%, with 12% decrease on queue makespan while still meeting all deadlines for HPC jobs.

sted, utgiver, år, opplag, sider
Los Alamitos: IEEE Computer Society, 2019
Emneord
Resource Management, High Performance Computing, Data Intensive Computing, Mesos, Slurm, Boostrapping
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
urn:nbn:se:umu:diva-155619 (URN)10.1109/CCGRID.2019.00054 (DOI)000483058700045 ()2-s2.0-85069469164 (Scopus ID)978-1-7281-0913-8 (ISBN)978-1-7281-0912-1 (ISBN)
Konferanse
19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGrid 2019, May 14-17, 2019, Larnaca, Cyprus
Merknad

Originally included in thesis in manuscript form

Tilgjengelig fra: 2019-01-24 Laget: 2019-01-24 Sist oppdatert: 2020-06-26bibliografisk kontrollert
Souza, A. P. (2018). Application-aware resource management for datacenters. (Licentiate dissertation). Umeå: Department of computing science, Umeå university
Åpne denne publikasjonen i ny fane eller vindu >>Application-aware resource management for datacenters
2018 (engelsk)Licentiatavhandling, med artikler (Annet vitenskapelig)
Alternativ tittel[sv]
Applikationsmedveten resurshantering för datacenter
Abstract [en]

High Performance Computing (HPC) and Cloud Computing datacenters are extensively used to steer and solve complex problems in science, engineering, and business, such as calculating correlations and making predictions. Already in a single datacenter server, there are thousands of hardware and software metrics – Key Performance Indicators (KPIs) – that individually and aggregated can give insight in the performance, robustness, and efficiency of the datacenter and the provisioned applications. At the datacenter level, the number of KPIs is even higher. The fast growing interest on datacenter management from both public and industry together with the rapid expansion in scale and complexity of datacenter resources and the services being provided on them have made monitoring, profiling, controlling, and provisioning compute resources dynamically at runtime into a challenging and complex task. Commonly, correlations of application KPIs, like response time and throughput, with resource capacities show that runtime systems (e.g., containers or virtual machines) that are used to provision these applications do not utilize available resources efficiently. This reduces datacenter efficiency, which in term results in higher operational costs and longer waiting times for results.

The goal of this thesis is to develop tools and autonomic techniques for improving datacenter operations, management and utilization, while improving and/or minimizing impacts on applications performance. To this end, we make use of application resource descriptors to create a library that dynamically adjusts the amount of resources used, enabling elasticity for scientific workflows in HPC datacenters. For mission critical applications, high availability is of great concern since these services must be kept running even in the event of system failures. By modeling and correlating specific resource counters, like CPU, memory and network utilization, with the number of runtime synchronizations, we present adaptive mechanisms to dynamically select which fault tolerant mechanism to use. Likewise, for scientific applications we propose a hybrid extensible architecture for dual-level scheduling of data intensive jobs in HPC infrastructures, allowing operational simplification, on-boarding of new types of applications and achieving greater job throughput with higher overall datacenter efficiency.

sted, utgiver, år, opplag, sider
Umeå: Department of computing science, Umeå university, 2018. s. 28
Serie
Report / UMINF, ISSN 0348-0542 ; 18.14
Emneord
Resource Management, High Performance Computing, Cloud Computing
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
urn:nbn:se:umu:diva-155620 (URN)978-91-7601-971-9 (ISBN)
Presentation
2018-12-12, MA121, MIT-Huset, Umeå, 20:31 (engelsk)
Opponent
Veileder
Tilgjengelig fra: 2019-01-25 Laget: 2019-01-24 Sist oppdatert: 2020-09-14bibliografisk kontrollert
Souza, A., Papadopoulos, A. V., Tomás Bolivar, L., Gilbert, D. & Tordsson, J. (2018). Hybrid Adaptive Checkpointing for Virtual Machine Fault Tolerance. In: Li J., Chandra A., Guo T., Cai Y. (Ed.), Proceedings - 2018 IEEE International Conference on Cloud Engineering, IC2E 2018: . Paper presented at 2018 IEEE International Conference on Cloud Engineering (IC2E 2018), 17–20 April 2018, Orlando, Florida, USA (pp. 12-22). Institute of Electrical and Electronics Engineers (IEEE)
Åpne denne publikasjonen i ny fane eller vindu >>Hybrid Adaptive Checkpointing for Virtual Machine Fault Tolerance
Vise andre…
2018 (engelsk)Inngår i: Proceedings - 2018 IEEE International Conference on Cloud Engineering, IC2E 2018 / [ed] Li J., Chandra A., Guo T., Cai Y., Institute of Electrical and Electronics Engineers (IEEE), 2018, s. 12-22Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Active Virtual Machine (VM) replication is an application independent and cost-efficient mechanism for high availability and fault tolerance, with several recently proposed implementations based on checkpointing. However, these methods may suffer from large impacts on application latency, excessive resource usage overheads, and/or unpredictable behavior for varying workloads. To address these problems, we propose a hybrid approach through a Proportional-Integral (PI) controller to dynamically switch between periodic and on-demand check-pointing. Our mechanism automatically selects the method that minimizes application downtime by adapting itself to changes in workload characteristics. The implementation is based on modifications to QEMU, LibVirt, and OpenStack, to seamlessly provide fault tolerant VM provisioning and to enable the controller to dynamically select the best checkpointing mode. Our evaluation is based on experiments with a video streaming application, an e-commerce benchmark, and a software development tool. The experiments demonstrate that our adaptive hybrid approach improves both application availability and resource usage compared to static selection of a checkpointing method, with application performance gains and neglectable overheads.

sted, utgiver, år, opplag, sider
Institute of Electrical and Electronics Engineers (IEEE), 2018
Emneord
Fault Tolerance, Resource Management, Checkpoint, COLO, Control Theory
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
urn:nbn:se:umu:diva-152033 (URN)10.1109/IC2E.2018.00023 (DOI)000759774400002 ()2-s2.0-85048315473 (Scopus ID)978-1-5386-5009-7 (ISBN)978-1-5386-5008-0 (ISBN)
Konferanse
2018 IEEE International Conference on Cloud Engineering (IC2E 2018), 17–20 April 2018, Orlando, Florida, USA
Tilgjengelig fra: 2018-09-24 Laget: 2018-09-24 Sist oppdatert: 2023-09-05bibliografisk kontrollert
Organisasjoner
Identifikatorer
ORCID-id: ORCID iD iconorcid.org/0000-0001-6952-1195