- 专利标题: Preprocessing of string inputs in natural language processing
-
申请号: US15376923申请日: 2016-12-13
-
公开(公告)号: US10372816B2公开(公告)日: 2019-08-06
- 发明人: Charles E. Beller , Chengmin Ding , Allen Ginsberg , Elinna Shek
- 申请人: International Business Machines Corporation
- 申请人地址: US NY Armonk
- 专利权人: International Business Machines Corporation
- 当前专利权人: International Business Machines Corporation
- 当前专利权人地址: US NY Armonk
- 代理机构: Lieberman & Brandsdorfer, LLC
- 主分类号: G06F17/27
- IPC分类号: G06F17/27
摘要:
Natural language processing of raw text data for optimal sentence boundary placement. Raw text is extracted from a document and subject to cleaning. The extracted raw text is examined to identify preliminary sentence boundaries, which are used to identify potential sentences in the raw text. One or more potential sentences are assigned a well-formedness score. A value of the score correlates to whether the potential sentence is a truncated/ill-formed sentence or a well-formed sentence. One or more preliminary sentence boundaries are optimized depending on the value of the score of the potential sentence(s). Accordingly, the processing herein is an optimization that creates a sentence boundary optimized output.
公开/授权文献
信息查询