PRE-TRAINING TECHNIQUES FOR ENTITY EXTRACTION IN LOW RESOURCE DOMAINS
Abstract:
Embodiments of the present invention provide systems, methods, and computer storage media for pre-training entity extraction models to facilitate domain adaptation in resource-constrained domains. In an example embodiment, a first machine learning model is used to encode sentences of a source domain corpus and a target domain corpus into sentence embeddings. The sentence embeddings of the target domain corpus are combined into a target corpus embedding. Training sentences from the source domain corpus within a threshold of similarity to the target corpus embedding are selected. A second machine learning model is trained on the training sentences selected from the source domain corpus.
Public/Granted literature
Information query
Patent Agency Ranking
0/0