Systems and methods for vision-language distribution alignment

Invention Grant

US12112523B2 Systems and methods for vision-language distribution alignment 有权

Please log in to see more content

Patent Title: Systems and methods for vision-language distribution alignment
Application No.: US17589725

Application Date: 2022-01-31
Publication No.: US12112523B2

Publication Date: 2024-10-08
Inventor: Shu Zhang , Junnan Li , Ran Xu , Caiming Xiong , Chetan Ramaiah
Applicant: Salesforce, Inc.
Applicant Address: US CA San Francisco
Assignee: Salesforce, Inc.
Current Assignee: Salesforce, Inc.
Current Assignee Address: US CA San Francisco
Agency: Haynes and Boone, LLP
Main IPC: G06V10/776
IPC: G06V10/776 ; G06F16/56 ; G06F16/583 ; G06F40/126 ; G06F40/166 ; G06F40/284 ; G06V10/74 ; G06V10/80

Systems and methods for vision-language distribution alignment

Abstract:

Embodiments described herein a CROss-Modal Distribution Alignment (CROMDA) model for vision-language pretraining, which can be used for retrieval downstream tasks. In the CROMDA mode, global cross-modal representations are aligned on each unimodality. Specifically, a uni-modal global similarity between an image/text and the image/text feature queue are computed. A softmax-normalized distribution is then generated based on the computed similarity. The distribution thus takes advantage of property of the global structure of the queue. CROMDA then aligns the two distributions and learns a modal invariant global representation. In this way, CROMDA is able to obtain invariant property in each modality, where images with similar text representations should be similar and vice versa.

Public/Granted literature

US20230162490A1 SYSTEMS AND METHODS FOR VISION-LANGUAGE DISTRIBUTION ALIGNMENT Public/Granted day:2023-05-25

Information query

Espacenet

IPC分类:

G	物理
G06	计算；推算或计数
G06V	图像或视频识别或理解
G06V10/00	图像或视频识别或理解的安排（图像或视频中的字符识别 G06V30/10）
G06V10/70	.使用模式识别或机器学习（光学模式识别或电子计算 G06V10/88）
G06V10/77	..处理特征空间中的图像或视频特征；使用数据集成或数据缩减，例如主成分分析 [PCA] 或独立成分分析 [ICA] 或自组织图 [SOM]；盲源分离
G06V10/776	...验证; 性能评估