Fall detection for older adults: A comparative study of CNN-based deep learning and ViT architectures
DOI:
https://doi.org/10.18687/LACCEI2026.1.1.1347Keywords:
fall detection, convolutional neural network, older adults, vision transformer, attention modulesAbstract
This paper presents a comparative study on fall detection in older adults using four pre-trained convolutional neural network models (VGG16, VGG19, ResNet50V2 and ResNet101V2), and Vision Transformer for Image Classification (ViT) model. Falls among older adults remain one of the leading cause of injury and reduced quality of life, and real-time detection systems can allow timely intervention to minimize harm. The proposed models are evaluated on a publicly available dataset composed of RGB images categorized into falls and non-falls. Each image is pre-processed through cropping, resizing to 128×128 and 224×224 pixels for CNN-based and ViT models, respectively, and Min-Max normalization.Transfer learning is applied to fine-tune the models using ImageNet-initialized weights, modifying the final layers to address the binary classification task. Models are trained and tested under consistent conditions, and performance is evaluated using accuracy, precision, recall, and F1 score metrics, supported by confusion matrices and ROC curves. Among the models, ViT model achievesthe highest classification accuracy (98%) and demonstrates a strong balance across performance metrics, particularly in detecting actual falls cases, which is critical in healthcare applications. This study confirms the effectiveness of ViT models for fall detection from single-camera input without the need for wearable sensors.Downloads
Published
2026-07-27
Issue
Section
Articles
Copyright
Copyright (c) 2026 LACCEI
License
This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
LACCEI retains copyright of all published articles under the terms of its copyright transfer agreement. As the copyright holder, LACCEI distributes the articles to the public under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).
How to Cite
Charco Aguirre, J. L., Yanza Montalván, A. O., Cruz Chóez, A. M., & Mendoza Morán, V. D. R. (2026). Fall detection for older adults: A comparative study of CNN-based deep learning and ViT architectures. LACCEI, 1(14). https://doi.org/10.18687/LACCEI2026.1.1.1347