-
公开(公告)号:US20220392242A1
公开(公告)日:2022-12-08
申请号:US17819838
申请日:2022-08-15
Abstract: A method for training a text positioning model includes: obtaining a sample image, where the sample image contains a sample text to be positioned and a text marking box for the sample text; inputting the sample image into a text positioning model to be trained to position the sample text, and outputting a prediction text box for the sample image; obtaining a sample prior anchor box corresponding to the sample image; and adjusting model parameters of the text positioning model based on the sample prior anchor box, the text marking box and the prediction text box, and continuing training the adjusted text positioning model based on a next sample image until model training is completed, to generate a target text positioning model.
-
公开(公告)号:US20220253631A1
公开(公告)日:2022-08-11
申请号:US17501221
申请日:2021-10-14
Inventor: Yulin LI , Ju HUANG , Qunyi XIE , Xiameng QIN , Chengquan ZHANG , Jingtuo LIU
Abstract: The present disclosure discloses an image processing method, an electronic device and a storage medium, and relates to the field of artificial intelligence technologies, and particularly to the fields of computer vision technologies, deep learning technologies, or the like. The image processing method includes: acquiring a multi-modal feature of each of at least one text region in an image, the multi-modal feature including features in plural dimensions; performing a global attention processing operation on the multi-modal feature of each text region to obtain a global attention feature of each text region; determining a category of each text region based on the global attention feature of each text region; and constructing structured information based on text content and the category of each text region.
-
公开(公告)号:US20210406468A1
公开(公告)日:2021-12-30
申请号:US17161466
申请日:2021-01-28
Inventor: Xiameng QIN , Yulin LI , Qunyi XIE , Ju HUANG , Junyu HAN
IPC: G06F40/279 , G06N3/08 , G06N3/04 , G06F16/532 , G06F16/583 , G06K9/20 , G06K9/62 , G06K9/46
Abstract: The present disclosure provides a method for visual question answering, which relates to a field of computer vision and natural language processing. The method includes: acquiring an input image and an input question; constructing a Visual Graph based on the input image, wherein the Visual Graph comprises a Node Feature and an Edge Feature; updating the Node Feature by using the Node Feature and the Edge Feature to obtain an updated Visual Graph; determining a question feature based on the input question; fusing the updated Visual Graph and the question feature to obtain a fused feature; and generating a predicted answer for the input image and the input question based on the fused feature. The present disclosure further provides an apparatus for visual question answering, a computer device and a non-transitory computer-readable storage medium.
-
公开(公告)号:US20210192696A1
公开(公告)日:2021-06-24
申请号:US17151783
申请日:2021-01-19
Inventor: Qunyi XIE , Xiameng QIN , Yulin LI , Junyu HAN , Shengxian ZHU
Abstract: Embodiments of the present disclosure provide a method and apparatus for correcting a distorted document image, where the method for correcting a distorted document image includes: obtaining a distorted document image; and inputting the distorted document image into a correction model, and obtaining a corrected image corresponding to the distorted document image; where the correction model is a model obtained by training with a set of image samples as inputs and a corrected image corresponding to each image sample in the set of image samples as an output, and the image samples are distorted. By inputting the distorted document image to be corrected into the correction model, the corrected image corresponding to the distorted document image can be obtained through the correction model, which realizes document image correction end-to-end, improves accuracy of the document image correction, and extends application scenarios of the document image correction.
-
5.
公开(公告)号:US20240021000A1
公开(公告)日:2024-01-18
申请号:US18113178
申请日:2023-02-23
Inventor: Xiameng QIN , Yulin LI , Xiaoqiang ZHANG , Ju HUANG , Qunyi XIE , Kun YAO
IPC: G06V30/19 , G06V30/148
CPC classification number: G06V30/1918 , G06V30/15 , G06V30/19127 , G06V30/19147
Abstract: There is provided an image-based information extraction model, method, and apparatus, a device, and a storage medium, which relates to the field of artificial intelligence (AI) technologies, specifically to fields of deep learning, image processing, computer vision technologies, and is applicable to optical character recognition (OCR) and other scenarios. A specific implementation solution involves: acquiring a to-be-extracted first image and a category of to-be-extracted information; and inputting the first image and the category into a pre-trained information extraction model to perform information extraction on the first image to obtain text information corresponding to the category.
-
公开(公告)号:US20220301334A1
公开(公告)日:2022-09-22
申请号:US17832735
申请日:2022-06-06
Inventor: Yuechen YU , Yulin LI , Chengquan ZHANG , Kun YAO
IPC: G06V30/416 , G06F40/18 , G06V30/413
Abstract: The present disclosure provides a table generating method and apparatus, an electronic device, a storage medium and a product. A specific implementation is: recognizing at least one table object in a to-be-recognized image and obtaining a table property respectively corresponding to the at least one table object, where the table property of any table object includes a cell property or a non-cell property; determining at least one target object with the cell property in the at least one table object; determining a cell region respectively corresponding to the at least one target object to obtain cell position information respectively corresponding to the at least one target object; generating a spreadsheet corresponding to the to-be-recognized image according to the cell position information respectively corresponding to the at least one target object.
-
公开(公告)号:US20220027611A1
公开(公告)日:2022-01-27
申请号:US17498226
申请日:2021-10-11
Inventor: Yuechen YU , Chengquan ZHANG , Yulin LI , Xiaoqiang ZHANG , Ju HUANG , Xiameng QIN , Kun YAO , Jingtuo LIU , Junyu HAN , Errui DING
Abstract: Provided are an image classification method and apparatus, an electronic device and a storage medium, relating to the field of artificial intelligence and, in particular, to computer vision and deep learning. The method includes inputting a to-be-classified document image into a pretrained neural network and obtaining a feature submap of each text box of the to-be-classified document image by use of the neural network; inputting the feature submap of each text box, a semantic feature corresponding to preobtained text information of each text box and a position feature corresponding to preobtained position information of each text box into a pretrained multimodal feature fusion model and fusing, by use of the multimodal feature fusion model, the three into a multimodal feature corresponding to each text box; and classifying the to-be-classified document image based on the multimodal feature corresponding to each text box.
-
公开(公告)号:US20210406592A1
公开(公告)日:2021-12-30
申请号:US17182987
申请日:2021-02-23
Inventor: Yulin LI , Xiameng QIN , Ju HUANG , Qunyi XIE , Junyu HAN
IPC: G06K9/62 , G06K9/46 , G06F40/279
Abstract: The present disclosure provides a method for visual question answering. The method includes: acquiring an input image and an input question; constructing a visual graph based on the input image, wherein the visual graph comprises a first node feature and a first edge feature; constructing a question graph based on the input question, wherein the question graph comprises a second node feature and a second edge feature; performing a multimodal fusion on the visual graph and the question graph to obtain an updated visual graph and an updated question graph; determining a question feature based on the input question; determining a fusion feature based on the updated visual graph, the updated question graph and the question feature; and generating a predicted answer for the input image and the input question. The present disclosure further provides an apparatus for visual question answering, a computer device and a medium.
-
公开(公告)号:US20210390294A1
公开(公告)日:2021-12-16
申请号:US17139403
申请日:2020-12-31
Inventor: Xiangkai Huang , Qiaoyi LI , Yulin LI , Ju Huang , Duohao Qin , Xiameng Qin , Minghao Liu , Junyu Han , Jiangliang Guo
Abstract: Embodiments of the present disclosure disclose an image table extraction method and apparatus, an electronic device, a storage media, and a training method for a table extraction model, which relate to the field of artificial intelligence technologies and cloud computing technologies, including: acquiring an image to be processed;
generating a table of the image to be processed according to a table extraction model, where the table extraction model is obtained according to a field position feature, an image feature, and a text feature of a sample image; and filling text information of the image to be processed into the table.
-
-
-
-
-
-
-
-