Effect of Hyper-Threading in Latency-Critical Multithreaded Cloud Applications and Utilization Analysis of the Major System Resources

[EN] Multithreaded latency-critical applications represent an important subset of workloads running on public cloud systems. Most of these systems deploy powerful computing servers including Intel Hyper-Threading processors. Understanding how performance is affected by the consumption of the main sy...

ver descrição completa

Detalhes bibliográficos
Autores: Pons-Escat, Lucía|||0000-0002-4582-7744, Feliu-Pérez, Josué|||0000-0003-3017-4266, Petit Martí, Salvador Vicente|||0000-0003-2426-4134, Pons Terol, Julio|||0000-0002-5654-6753, Gómez Requena, María Engracia|||0000-0003-1466-4118, Sahuquillo Borrás, Julio|||0000-0001-8630-4846, Puche-Lara, José, Huang, Chaoyi
Formato: artículo
Fecha de publicación:2022
País:España
Recursos:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/197359
Acesso em linha:https://riunet.upv.es/handle/10251/197359
Access Level:acceso abierto
Palavra-chave:Cloud computing
Latency-critical workloads
Level of load
Tail latency
Resource sharing
Hyper-Threading
ARQUITECTURA Y TECNOLOGIA DE COMPUTADORES
Descrição
Resumo:[EN] Multithreaded latency-critical applications represent an important subset of workloads running on public cloud systems. Most of these systems deploy powerful computing servers including Intel Hyper-Threading processors. Understanding how performance is affected by the consumption of the main system resources is a major concern for cloud providers in order to devise virtualization strategies that improve the system efficiency. With this aim, this paper first characterizes the impact of QPS on tail latency, analyzing different scenarios varying the number of threads and the thread-to-core allocation (single-task and multi-task execution) policy. The characterization study reveals that the performance of some applications does not scale with the number of threads, and the performance of some others is insensitive to the Hyper-Threading technology, so they can be allocated in less physical cores and improve system utilization. Identifying these applications, however, at run-time is challenging. Despite identifying these applications at run-time is challenging, this paper shows that they can be successfully detected at run-time by analyzing the utilization trend of the major system resources. In addition to CPU, we have also studied how assigning the share of each application of other major shared system resources impacts on performance. We outline considerations cloud providers should take into account to improve performance and resource utilization.