Invention Application
- Patent Title: MODALITY ADAPTIVE INFORMATION RETRIEVAL
-
Application No.: US17153130Application Date: 2021-01-20
-
Publication No.: US20220230061A1Publication Date: 2022-07-21
- Inventor: Hrituraj Singh , Jatin Lamba , Denil Pareshbhai Mehta , Balaji Vasan Srinivasan , Anshul Nasery , Aishwarya Agarwal
- Applicant: Adobe Inc.
- Applicant Address: US CA San Jose
- Assignee: Adobe Inc.
- Current Assignee: Adobe Inc.
- Current Assignee Address: US CA San Jose
- Main IPC: G06N3/08
- IPC: G06N3/08 ; G06F16/242 ; G06N3/04 ; G06K9/62 ; G06K9/00 ; G06F40/20

Abstract:
In some embodiments, a multimodal computing system receives a query and identifies, from source documents, text passages and images that are relevant to the query. The multimodal computing system accesses a multimodal question-answering model that includes a textual stream of language models and a visual stream of language models. Each of the textual stream and the visual stream contains a set of transformer-based models and each transformer-based model includes a cross-attention layer using data generated by both the textual stream and visual stream of language models as an input. The multimodal computing system identifies text relevant to the query by applying the textual stream to the text passages and computes, using the visual stream, relevance scores of the images to the query, respectively. The multimodal computing system further generates a response to the query by including the text and/or an image according to the relevance scores.
Public/Granted literature
- US12198048B2 Modality adaptive information retrieval Public/Granted day:2025-01-14
Information query