-
公开(公告)号:US20240062560A1
公开(公告)日:2024-02-22
申请号:US17901617
申请日:2022-09-01
Applicant: Google LLC
Inventor: Shangbang Long , Siyang Qin , Dmitry Panteleev , Alessandro Bissacco , Yasuhisa Fujii , Michail Raptis
IPC: G06V20/62 , G06V30/414 , G06V10/82 , G06V30/14
CPC classification number: G06V20/63 , G06V30/414 , G06V10/82 , G06V30/1448
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for jointly performing text detection and layout analysis. In one aspect, a method comprises processing the image and a set of object queries to generate an encoded representation of the image and an encoded representation of the set of object queries; processing the encoded representation of the image and the encoded representation of the set of object queries to generate a set of text detection masks; processing the encoded representation of the set of object queries to generate layout relevance measures; processing the encoded representation of the set of object queries to generate textness scores for the text detection masks; generating a text detection output that defines respective areas of the image that include text items; and generating a layout analysis output that defines clusters of respective areas of the image identified by the text detection masks.
-
公开(公告)号:US20250022301A1
公开(公告)日:2025-01-16
申请号:US18772414
申请日:2024-07-15
Applicant: Google LLC
Inventor: Shangbang Long , Siyang Qin , Yasuhisa Fujii , Alessandro Bissacco , Michail Raptis
IPC: G06V30/20 , G06V30/14 , G06V30/166 , G06V30/18 , G06V30/19
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for detecting text instances of arbitrary shapes, sizes, and locations. In one aspect, a method comprises processing an image depicting one or more text instances, generating a respective prediction for each character in a sequence of characters that are predicted to be depicted in the text instance, the respective prediction comprising a respective character class to which the predicted character belongs, the respective character class selected from a set that includes printable character classes and a space character class and a bounding box that contains the character within the image, and grouping the sequence of characters into a plurality of words based on locations of characters that are predicted to belong to the space character class.
-