TEXT-CONDITIONED VIDEO REPRESENTATION
    1.
    发明公开

    公开(公告)号:US20230351753A1

    公开(公告)日:2023-11-02

    申请号:US17894738

    申请日:2022-08-24

    CPC classification number: G06V20/47 G06V20/41

    Abstract: A text-video recommendation model determines relevance of a text to a video in a text-video pair (e.g., as a relevance score) with a text embedding and a text-conditioned video embedding. The text-conditioned video embedding is a representation of the video used for evaluating the relevance of the video to the text, where the representation itself is a function of the text it is evaluated for. As such, the input text may be used to weigh or attend to different frames of the video in determining the text-conditioned video embedding. The representation of the video may thus differ for different input texts for comparison. The text-conditioned video embedding may be determined in various ways, such as with a set of the most-similar frames to the input text (the top-k frames) or may be based on an attention function based on query, key, and value projections.

Patent Agency Ranking