NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps

Convolutional neural networks (CNNs) have become the dominant neural network architecture for solving many stateof- the-art (SOA) visual processing tasks. Even though Graphical Processing Units (GPUs) are most often used in training and deploying CNNs, their power efficiency is less than 10 GOp/s/W...

ver descrição completa

Detalhes bibliográficos
Autores: Aimar, Alessandro, Mostafa, Hesham, Calabrese, Enrico, Ríos Navarro, José Antonio, Tapiador Morales, Ricardo, Lungu, Iulia-Alexandra, Milde, Moritz B., Corradi, Federico, Linares Barranco, Alejandro, Liu, Shih-Chii, Delbruck, Tobi
Tipo de documento: artigo
Estado:Versión enviada para evaluación y publicación
Data de publicação:2019
País:España
Recursos:Universidad de Sevilla (US)
Repositório:idUS. Depósito de Investigación de la Universidad de Sevilla
OAI Identifier:oai:idus.us.es:11441/92660
Acesso em linha:https://hdl.handle.net/11441/92660
https://doi.org/10.1109/TNNLS.2018.2852335
Access Level:Acceso aberto
Palavra-chave:Convolutional Neural Networks (CNN)
VLSI
FPGA
Computer vision
Artificial intelligence
id ES_808e2340660cab8dc4a4f457ce8eed65
oai_identifier_str oai:idus.us.es:11441/92660
network_acronym_str ES
network_name_str España
repository_id_str
spelling NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature MapsAimar, AlessandroMostafa, HeshamCalabrese, EnricoRíos Navarro, José AntonioTapiador Morales, RicardoLungu, Iulia-AlexandraMilde, Moritz B.Corradi, FedericoLinares Barranco, AlejandroLiu, Shih-ChiiDelbruck, TobiConvolutional Neural Networks (CNN)VLSIFPGAComputer visionArtificial intelligenceConvolutional neural networks (CNNs) have become the dominant neural network architecture for solving many stateof- the-art (SOA) visual processing tasks. Even though Graphical Processing Units (GPUs) are most often used in training and deploying CNNs, their power efficiency is less than 10 GOp/s/W for single-frame runtime inference.We propose a flexible and efficient CNN accelerator architecture called NullHop that implements SOA CNNs useful for low-power and low-latency application scenarios. NullHop exploits the sparsity of neuron activations in CNNs to accelerate the computation and reduce memory requirements. The flexible architecture allows high utilization of available computing resources across kernel sizes ranging from 1x1 to 7x7. NullHop can process up to 128 input and 128 output feature maps per layer in a single pass. We implemented the proposed architecture on a Xilinx Zynq FPGA platform and present results showing how our implementation reduces external memory transfers and compute time in five different CNNs ranging from small ones up to the widely known large VGG16 and VGG19 CNNs. Post-synthesis simulations using Mentor Modelsim in a 28nm process with a clock frequency of 500MHz show that the VGG19 network achieves over 450GOp/s. By exploiting sparsity, NullHop achieves an efficiency of 368%, maintains over 98% utilization of the MAC units, and achieves a power efficiency of over 3TOp/s/W in a core area of 6.3mm2. As further proof of NullHop’s usability, we interfaced its FPGA implementation with a neuromorphic event camera for real time interactive demonstrations.IEEE Computer SocietyArquitectura y Tecnología de ComputadoresTEP-108: Robótica y Tecnología de Computadores Aplicada a la Rehabilitación2019info:eu-repo/semantics/articleinfo:eu-repo/semantics/submittedVersionapplication/pdfapplication/pdfhttps://hdl.handle.net/11441/92660https://doi.org/10.1109/TNNLS.2018.2852335reponame:idUS. Depósito de Investigación de la Universidad de Sevillainstname:Universidad de Sevilla (US)InglésIEEE Transactions on Neural Networks and Learning Systems, 30 (3), 644-656.https://ieeexplore.ieee.org/document/8421093info:eu-repo/semantics/openAccessoai:idus.us.es:11441/926602026-06-17T12:51:07Z
dc.title.none.fl_str_mv NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
title NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
spellingShingle NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
Aimar, Alessandro
Convolutional Neural Networks (CNN)
VLSI
FPGA
Computer vision
Artificial intelligence
title_short NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
title_full NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
title_fullStr NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
title_full_unstemmed NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
title_sort NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
dc.creator.none.fl_str_mv Aimar, Alessandro
Mostafa, Hesham
Calabrese, Enrico
Ríos Navarro, José Antonio
Tapiador Morales, Ricardo
Lungu, Iulia-Alexandra
Milde, Moritz B.
Corradi, Federico
Linares Barranco, Alejandro
Liu, Shih-Chii
Delbruck, Tobi
author Aimar, Alessandro
author_facet Aimar, Alessandro
Mostafa, Hesham
Calabrese, Enrico
Ríos Navarro, José Antonio
Tapiador Morales, Ricardo
Lungu, Iulia-Alexandra
Milde, Moritz B.
Corradi, Federico
Linares Barranco, Alejandro
Liu, Shih-Chii
Delbruck, Tobi
author_role author
author2 Mostafa, Hesham
Calabrese, Enrico
Ríos Navarro, José Antonio
Tapiador Morales, Ricardo
Lungu, Iulia-Alexandra
Milde, Moritz B.
Corradi, Federico
Linares Barranco, Alejandro
Liu, Shih-Chii
Delbruck, Tobi
author2_role author
author
author
author
author
author
author
author
author
author
dc.contributor.none.fl_str_mv Arquitectura y Tecnología de Computadores
TEP-108: Robótica y Tecnología de Computadores Aplicada a la Rehabilitación
dc.subject.none.fl_str_mv Convolutional Neural Networks (CNN)
VLSI
FPGA
Computer vision
Artificial intelligence
topic Convolutional Neural Networks (CNN)
VLSI
FPGA
Computer vision
Artificial intelligence
description Convolutional neural networks (CNNs) have become the dominant neural network architecture for solving many stateof- the-art (SOA) visual processing tasks. Even though Graphical Processing Units (GPUs) are most often used in training and deploying CNNs, their power efficiency is less than 10 GOp/s/W for single-frame runtime inference.We propose a flexible and efficient CNN accelerator architecture called NullHop that implements SOA CNNs useful for low-power and low-latency application scenarios. NullHop exploits the sparsity of neuron activations in CNNs to accelerate the computation and reduce memory requirements. The flexible architecture allows high utilization of available computing resources across kernel sizes ranging from 1x1 to 7x7. NullHop can process up to 128 input and 128 output feature maps per layer in a single pass. We implemented the proposed architecture on a Xilinx Zynq FPGA platform and present results showing how our implementation reduces external memory transfers and compute time in five different CNNs ranging from small ones up to the widely known large VGG16 and VGG19 CNNs. Post-synthesis simulations using Mentor Modelsim in a 28nm process with a clock frequency of 500MHz show that the VGG19 network achieves over 450GOp/s. By exploiting sparsity, NullHop achieves an efficiency of 368%, maintains over 98% utilization of the MAC units, and achieves a power efficiency of over 3TOp/s/W in a core area of 6.3mm2. As further proof of NullHop’s usability, we interfaced its FPGA implementation with a neuromorphic event camera for real time interactive demonstrations.
publishDate 2019
dc.date.none.fl_str_mv 2019
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/submittedVersion
format article
status_str submittedVersion
dc.identifier.none.fl_str_mv https://hdl.handle.net/11441/92660
https://doi.org/10.1109/TNNLS.2018.2852335
url https://hdl.handle.net/11441/92660
https://doi.org/10.1109/TNNLS.2018.2852335
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv IEEE Transactions on Neural Networks and Learning Systems, 30 (3), 644-656.
https://ieeexplore.ieee.org/document/8421093
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv IEEE Computer Society
publisher.none.fl_str_mv IEEE Computer Society
dc.source.none.fl_str_mv reponame:idUS. Depósito de Investigación de la Universidad de Sevilla
instname:Universidad de Sevilla (US)
instname_str Universidad de Sevilla (US)
reponame_str idUS. Depósito de Investigación de la Universidad de Sevilla
collection idUS. Depósito de Investigación de la Universidad de Sevilla
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869411907281092608
score 15.198674