Self-Supervised Learning via Multi-Transformation Classification for Action Recognition

Duc Quang Vu, Ngan Le, Jia Ching Wang

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

2 Scopus citations

Abstract

Self-supervised tasks have been utilized to build useful representations that can be used in downstream tasks when the annotation is unavailable. In this paper, we introduce a self-supervised video representation learning method based on the multi-transformation classification to efficiently classify human actions. Self-supervised learning on various transformations not only provides richer contextual information but also enables the visual representation more robust to the transforms. The spatio-temporal representation of the video is learned in a self-supervised manner by classifying seven different transformations i.e. rotation, clip inversion, permutation, split, join transformation, color switch, frame replacement, and noise addition. First, seven different video transformations are applied to video clips. Then the 3D convolutional neural networks are utilized to extract features for clips and these features are processed to classify the pseudo-labels. We use the learned models in pretext tasks as the pre-trained models and fine-tune them to recognize human actions in the downstream task. We have conducted the experiments on UCF101 and HMDB51 datasets together with C3D and 3D Resnet-18 as backbone networks. The experimental results have shown that our proposed framework outperformed other SOTA self-supervised action recognition approaches.

Original languageEnglish
Title of host publication2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798350379815
DOIs
StatePublished - 2024
Event2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024 - Niagara Falls, Canada
Duration: 15 Jul 202419 Jul 2024

Publication series

Name2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024

Conference

Conference2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024
Country/TerritoryCanada
CityNiagara Falls
Period15/07/2419/07/24

Keywords

  • 3D ResNet
  • Action Recognition
  • C3D
  • multi-transformation
  • Self-supervised learning

Fingerprint

Dive into the research topics of 'Self-Supervised Learning via Multi-Transformation Classification for Action Recognition'. Together they form a unique fingerprint.

Cite this