Is There an Optimal Depth of Residual Networks?
Journal
Communications in Computer and Information Science
ISBN
978-3-031-87569-4
Type
conference paper
Date Issued
2025-05-03
Author(s)
Abstract
Although deep neural networks have given the name to the domain of "deep learning", it is still an unresolved question of which depth is the best choice for a given task. A large depth brings problems with the convergence of gradientbased learning algorithms through the phenomenon of vanishing gradient. On the other hand, it is a widespread opinion that a small depth has insufficient representational power for many tasks. The discovery of the concept of residual connections-an identity mapping parallel to a conventional layer-has alleviated the convergence problem so that the discussion of optimum depth lost a part of its motivation, resulting in the assumption of "the more the better". The work presented here shows that a shallow architecture of parallel layers has comparable expressive power as a deep stack of residual layers. This is theoretically justified by expanding the residual layer stack analogical to the Taylor expansion, truncating the higher-order terms into a single broad layer composed of original layers in parallel. This hypothesis has been confirmed by computing experiments with the widespread computer vision benchmark datasets MNIST and CIFAR-10. The 6,912 runs have shown that the shallow and the deep architectures do not substantially differ in performance on both training and validation sets if the total number of parameters is equal. The rough equivalence of the two extreme (deep and shallow) architectures suggests the possibility that an intermediary architecture may be superior. Another series of computing experiments disclosed that the performance does not substantially differ even then. The conclusion is that the performance of an architecture depends more substantially on the total number of parameters than on the sequential or parallel connection of layers.
Language
English
Keywords
Deep learning
Residual networks
Expressive power
Computer vision
Convolutional neural networks
HSG Classification
contribution to scientific community
Refereed
Yes
Book title
Knowledge Discovery, Knowledge Engineering and Knowledge Management
Publisher
Springer Nature (Switzerland)
Publisher place
Cham
Volume
2454
Start page
91
End page
111
Contact Email Address
bernhard.bermeitinger@unisg.ch
File(s)![Thumbnail Image]()
Name
978-3-031-87569-4_5.pdf
Type
Main Article
Description
Published article
Size
983.25 KB
Format
Adobe PDF
Checksum (MD5)
d544bfcb81c77626075dfc9883fccc1d