Improved Hand Gesture Recognition in Uncontrolled Scenario Using CNN Integrated with Spatial Attention Mechanism
Résumé fourni par la source
This study introduces an advanced approach to hand gesture recognition that uses Convolutional Neural Networks (CNNs) augmented with a spatial attention mechanism. The core aim is to improve the model’s capacity to identify detailed spatialtemporal patterns within dynamic hand gestures, particularly under uncontrolled environmental conditions. The proposed method processes both RGB video frames and optical flow data to accurately recognize a range of static and dynamic gestures. By incorporating spatial attention into the CNN framework, the model enhances its ability to extract relevant features and prioritize crucial gesture elements. The performance is validated using the Leave-One-Person-Out (LOPO) cross-validation method, ensuring strong generalizability across different individuals. Results show that the model achieves an average recognition accuracy of 91.64%, with precision, recall, and F1-score all surpassing 90%. The inclusion of attention mechanisms leads to more than a 1% performance boost over traditional CNNs without attention, especially in challenging scenarios with inconsistent lighting and complex backgrounds. When compared to baseline models, the proposed method significantly outperforms conventional CNN and LSTM architectures. These outcomes highlight the model’s potential for impactful use in humancomputer interaction, prosthetic control, and assistive technology applications.