Web: http://arxiv.org/abs/2201.12596

Sept. 15, 2022, 1:13 a.m. | Zejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen, Zhongyu Wei, Xuanjing Huang

cs.CV updates on arXiv.org arxiv.org

Previous vision-language pre-training models mainly construct multi-modal
inputs with tokens and objects (pixels) followed by performing cross-modality
interaction between them. We argue that the input of only tokens and object
features limits high-level semantic alignment like phrase-to-region grounding.
Meanwhile, multi-level alignments are inherently consistent and able to
facilitate the representation learning synergistically. Therefore, in this
paper, we propose to learn Multi-level semantic alignment for Vision-language
Pre-TRaining (MVPTR). In MVPTR, we follow the nested structure of both
modalities to introduce concepts …

alignment arxiv language pre-training semantic stage training vision

More from arxiv.org / cs.CV updates on arXiv.org

Postdoctoral Fellow: ML for autonomous materials discovery

@ Lawrence Berkeley National Lab | Berkeley, CA

Research Scientists

@ ODU Research Foundation | Norfolk, Virginia

Embedded Systems Engineer (Robotics)

@ Neo Cybernetica | Bedford, New Hampshire

2023 Luis J. Alvarez and Admiral Grace M. Hopper Postdoc Fellowship in Computing Sciences

@ Lawrence Berkeley National Lab | San Francisco, CA

Senior Manager Data Scientist

@ NAV | Remote, US

Senior AI Research Scientist

@ Earth Species Project | Remote anywhere

Research Fellow- Center for Security and Emerging Technology (Multiple Opportunities)

@ University of California Davis | Washington, DC

Staff Fellow - Data Scientist

@ U.S. FDA/Center for Devices and Radiological Health | Silver Spring, Maryland

Staff Fellow - Senior Data Engineer

@ U.S. FDA/Center for Devices and Radiological Health | Silver Spring, Maryland

Research Engineer - VFX, Neural Compositing

@ Flawless | Los Angeles, California, United States

[Job-TB] Senior Data Engineer

@ CI&T | Brazil

Data Analytics Engineer

@ The Fork | Paris, France