摘要:
Disclosed is a method of segmenting touching numeral strings contained in handwritten touching numeral strings, and recognizing the numeral strings by use of feature information and recognized results provided by inherent structure of digits. The method comprises the steps of: receiving a handwritten numeral string extracted from a pattern document; smoothing a curved numeral image of the handwritten numeral string, and searching connecting components in the numeral image; determining whether or not the numeral string is a touching numeral string; if it is determined that the numeral string is the touching numeral string, searching a contour of the touching numeral string image; searching candidate segmentation points in the contour, and segmenting sub-images; computing a segmentation confidence value on each segmented sub-image by use of a segmentation error function to select the sub-image with the highest segmentation confidence value as a segmented numeral image in the touching numeral string image; if it is determined in the step c that the numeral string is not the touching numeral string, extracting a feature to recognize the segmented numeral image; segmenting the numeral image selected from the touching numeral string in the highest segmenting confidence value; and obtaining remaining numeral string image.
摘要:
A method for analyzing structure of a treatise type of document image in order to detect a title, an author and an abstract region and recognize the content in each of the regions is provided. In order to analyze the structure of a treatise type of document, first, the document image divided into a number of regions and the divided regions are classified into text regions and non-text regions according to attributes of the regions. And then, the candidate regions representing an abstract and an introduction is selected, thereafter word regions are extracted from the candidate regions, and an abstract content portion is determined. Thereafter, the title and the author are separated by using the basic form and the type definition representing an arrangement of each of journals. Finally, the content of the separated regions is recognized to generate said table of contents.
摘要:
Disclosed is a method for recognizing multi-language printed documents, a method for extracting character features according to the present invention, the method comprising the steps of: a) normalizing characters to a fixed size; b) converting the size-fixed characters into mesh-type characters; c) extracting stroke features of each of the mesh-type characters; d) extracting non-stroke features of each of the mesh-type characters; and e) extracting the character features using the stroke features and the non-stroke features. The present invention provides a high recognition rate irrespective of the size and modification of the characters, by extracting the character feature from the stroke and non-stroke in the mesh block.