April 20, 2022, 11:11 p.m. | /u/ephemeral_lives

Computer Vision www.reddit.com

Hello everyone!

For a project of mine involving video data (temporal RGB data), I was contemplating the usage of transformers/attention for the task head to perform a specific task, say action recognition in a video. The high-level idea was to use different (task-relevant) features and then use an attention mechanism to do the final task. I had a few questions for the same -

1. How feasible it would be to train such a network (transformer/attention part?)
2. If this …

computervision transformers videos

Senior Machine Learning Engineer (MLOps)

@ Promaton | Remote, Europe

Applied Scientist, Control Stack, AWS Center for Quantum Computing

@ Amazon.com | Pasadena, California, USA

Specialist Marketing with focus on ADAS/AD f/m/d

@ AVL | Graz, AT

Machine Learning Engineer, PhD Intern

@ Instacart | United States - Remote

Supervisor, Breast Imaging, Prostate Center, Ultrasound

@ University Health Network | Toronto, ON, Canada

Senior Manager of Data Science (Recommendation Science)

@ NBCUniversal | New York, NEW YORK, United States