Temporal Convolutional Networks: A Unified Approach to Action\n Segmentation
Le résumé fourni par la source
The dominant paradigm for video-based action segmentation is composed of two\nsteps: first, for each frame, compute low-level features using Dense\nTrajectories or a Convolutional Neural Network that encode spatiotemporal\ninformation locally, and second, input these features into a classifier that\ncaptures high-level temporal relationships, such as a Recurrent Neural Network\n(RNN). While often effective, this decoupling requires specifying two separate\nmodels, each with their own complexities, and prevents capturing more nuanced\nlong-range spatiotemporal relationships. We propose a unified approach, as\ndemonstrated by our Temporal Convolutional Network (TCN), that hierarchically\ncaptures relationships at low-, intermediate-, and high-level time-scales. Our\nmodel achieves superior or competitive performance using video or sensor data\non three public action segmentation datasets and can be trained in a fraction\nof the time it takes to train an RNN.\n
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.