Neural Network Training Techniques for Human Activity Recognition Using Visual Data
Date Issued
June 21, 2024
Type
Διδακτορική Διατριβή
Abstract
The goal of this Thesis is to investigate approaches and training strategies for training deep networks architectures with the goal of Human Activity Recognition (HAR). In the pursuit to enhance (HAR) systems, this Thesis begins with a comprehensive exploration of various recent research efforts on the field of HAR. The research efforts begin with the introduction of a novel approach for 2D representation of 3D motion of skeletal joints, based on well-known 2D spectral image transformations (DFT, FFT, DCT, DST), demonstrating its efficacy in HAR tasks. However, as this approach transitioned to cross-view setups, a notable decline in performance emerged, driving research towards a thorough examination, aiming to uncover nuanced details that could shed light on potential avenues for improvement.
Therefore, recognizing this challenge, research turned into devising an advanced data augmentation approach using artificial data which resulted upon applying geometric rotation transformation to real samples and also imposed a view alignment step. This way it was demonstrated that cross-view scenarios may also exhibit strong performance, comparable to other evaluation setups. Through meticulous experimentation, this performance drop was mitigated, showcasing the adaptability of the proposed approach across different viewing perspectives and reinforcing its relevance in real-world scenarios.
The next step of research was the investigation of domain adaptation techniques, where adversarial training was leveraged to bridge performance gaps. While these endeavors yielded promising results under specific circumstances, it was also acknowledged that heightened complexity was associated with this approach, rendering it impractical for real-life applications. Therefore, the next step was to explore fusion approaches that combined various motion representations, i.e., RGB and skeletal data. These fusion techniques not only surpassed the performance of state-of-the-art methods but also underscored the versatility and robustness of the proposed methodology in capturing nuanced human actions across diverse contexts.
Finally, this Thesis concludes with a novel occlusion-based data augmentation method. This approach not only surpassed the efficacy of the initial work but also contributed significantly to overall HAR performance enhancement. By addressing occlusion challenges inherent in real-world scenarios, this augmentation technique further validated the practical applicability of the research findings. In essence, the journey of this Thesis through experimentation, adaptation, fusion, and augmentation underscores a holistic approach to advancing HAR systems, paving the way for more robust and versatile human activity recognition technologies in diverse environments.
Subjects
