Tech Reports | Knowledge Media Institute

Tech Report

Semi-Automatic Population of Ontologies from Text

This paper describes a system for semi-automatic population of ontologies with instances from unstructured text. The system is based on supervised learning and therefore learns extraction rules from annotated text and then applies those rules on newly documents for ontology population. It is based on three componentes: Marmot, a natural language processor; Crystal, a dictionary induction tool; and Badger, an information extraction tool. The important part of the entire cycle is a user who accepts, rejects or modifies newly extracted and suggested instances to be populated. A description of experiments performed with text corpus consisting of 91 documents is given in turn. The results cover the paper and support a presented hypothesis of assigning a rule confi-dence value to each extraction rule to improve the performance.

ID: kmi-04-18

Date: 2004

Author(s): David Celjuska, Maria Vargas-Vera

Resources: