Repository logo
Research Outputs
Projects
People
Statistics
  1. Home
  2. HSG CRIS
  3. HSG Publications
  4. Reducing the Transformer Architecture to a Minimum
Details

Reducing the Transformer Architecture to a Minimum

Journal
Proceedings of the 16th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management
Type
conference paper
Date Issued
2024-11
Author(s)
Bernhard Bermeitinger  
;
Tomas Hrycej  
;
Massimo Pavone  
;
Julianus Kath
;
Siegfried Handschuh  
DOI
10.5220/0012891000003838
Abstract
Transformers are a widespread and successful model architecture, particularly in Natural Language Processing (NLP) and Computer Vision (CV). The essential innovation of this architecture is the Attention Mechanism, which solves the problem of extracting relevant context information from long sequences in NLP and realistic scenes in CV. A classical neural network component, a Multi-Layer Perceptron (MLP), complements the attention mechanism. Its necessity is frequently justified by its capability of modeling nonlinear relationships. However, the attention mechanism itself is nonlinear through its internal use of similarity measures. A possible hypothesis is that this nonlinearity is sufficient for modeling typical application problems. As the MLPs usually contain the most trainable parameters of the whole model, their omission would substantially reduce the parameter set size. Further components can also be reorganized to reduce the number of parameters. Under some conditions, query and key matrices can be collapsed into a single matrix of the same size. The same is true about value and projection matrices, which can also be omitted without eliminating the substance of the attention mechanism. Initially, the similarity measure was defined asymmetrically, with peculiar properties such as that a token is possibly dissimilar to itself. A possible symmetric definition requires only half of the parameters. All these parameter savings make sense only if the representational performance of the architecture is not significantly reduced. A comprehensive empirical proof for all important domains would be a huge task. We have laid the groundwork by testing widespread CV benchmarks: MNIST, CIFAR-10, and, with restrictions, ImageNet. The tests have shown that simplified transformer architectures (a) without MLP, (b) with collapsed matrices, and (c) symmetric similarity matrices exhibit similar performance as the original architecture, saving up to 90 % of parameters without hurting the classification performance.
Language
English
Keywords
Attention Mechanism
Transformers
Computer Vision
Model Reduction
Deep Neural Networks
HSG Classification
contribution to scientific community
Publisher
SCITEPRESS - Science and Technology Publications
Start page
234
End page
241
Pages
8
Event Location
Porto, Portugal
Event Date
November 2024
Official URL
https://www.scitepress.org/PublicationsDetail.aspx?ID=REzv8Neq7FY=&t=1
URL
https://www.alexandria.unisg.ch/handle/20.500.14171/121344
Contact Email Address
bernhard.bermeitinger@unisg.ch
File(s)
Thumbnail Image

restricted

Name

KDIR2024 - Reducing the Transformer Architecture to a Minimum.pdf

Description
published version (restricted access)
Size

219.82 KB

Format

Adobe PDF

Checksum (MD5)

295d8917028788577229d105dc2e3423

Thumbnail Image

open.access

Name

KDIR2024 - Reducing the Transformer Architecture to a Minimum - Poster.pdf

Description
poster
Size

201.56 KB

Format

Adobe PDF

Checksum (MD5)

ba63759e8d5560a8ae0b6618433119d3

Support
HSG researchers can find instructions here for adding or importing publications (DOI, ORCID). Please send questions to alexandria@unisg.ch

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify