TY - GEN
T1 - Shallow Enough? A Cross-Architecture Study of Ultra-Low-Depth Neural Networks for Edge Inference
AU - Zhou, Chengwei
AU - Yu, Haotian
AU - Yukawa, Shoma
AU - Najafi, Deniz
AU - Angizi, Shaahin
AU - Datta, Gourav
N1 - Publisher Copyright: © 2026 Copyright held by the owner/author(s).
PY - 2026/6/22
Y1 - 2026/6/22
N2 - Edge deployment imposes strict latency, memory, and energy constraints that scale directly with network depth, yet the question of which architectural paradigm offers the best accuracy-efficiency tradeoff at ultra-low depth (four to six layers) remains open. We present the first systematic comparison of convolutional, pure transformer, and hybrid CNN-transformer models under fixed shallow depth budgets. We develop an analytical framework characterizing the representational cost of local convolutional versus global attention-based feature mixing as a function of depth, and validate it with experiments on ImageNet-1K across architectures spanning MobileNetV2, DeiT, Swin, MobileViT, and EfficientFormer. Beyond FLOPs and accuracy, we report real-world latency and energy measurements on an NVIDIA Jetson Nano and examine deployment feasibility on Cortex-M class microcontrollers. Our experiments reveal how accuracy, latency, and memory scale across all three architecture families as depth decreases, providing practitioners with direct guidance on which model class to choose for a given depth and hardware budget.
AB - Edge deployment imposes strict latency, memory, and energy constraints that scale directly with network depth, yet the question of which architectural paradigm offers the best accuracy-efficiency tradeoff at ultra-low depth (four to six layers) remains open. We present the first systematic comparison of convolutional, pure transformer, and hybrid CNN-transformer models under fixed shallow depth budgets. We develop an analytical framework characterizing the representational cost of local convolutional versus global attention-based feature mixing as a function of depth, and validate it with experiments on ImageNet-1K across architectures spanning MobileNetV2, DeiT, Swin, MobileViT, and EfficientFormer. Beyond FLOPs and accuracy, we report real-world latency and energy measurements on an NVIDIA Jetson Nano and examine deployment feasibility on Cortex-M class microcontrollers. Our experiments reveal how accuracy, latency, and memory scale across all three architecture families as depth decreases, providing practitioners with direct guidance on which model class to choose for a given depth and hardware budget.
KW - CNN
KW - Depth Reduction
KW - Edge Inference
KW - Hybrid Architectures
KW - Vision Transformers
UR - https://www.scopus.com/pages/publications/105045049621
UR - https://www.scopus.com/pages/publications/105045049621#tab=citedBy
U2 - 10.1145/3787109.3816404
DO - 10.1145/3787109.3816404
M3 - Conference contribution
T3 - GLSVLSI 2026 - Proceedings of the Great Lakes Symposium on VLSI 2026
SP - 220
EP - 225
BT - GLSVLSI 2026 - Proceedings of the Great Lakes Symposium on VLSI 2026
A2 - Chen, Fan
A2 - Zhou, Peipei
A2 - Gu, Jie
A2 - Trivedi, Amit R.
A2 - Yang, Xiaoxuan
PB - Association for Computing Machinery, Inc
T2 - 36th Great Lakes Symposium on VLSI, GLSVLSI 2026
Y2 - 22 June 2026 through 24 June 2026
ER -