Large Dataset for Keyphrases Extraction

Krapivin, Mikalai and Autaeu, Aliaksandr and Marchese, Maurizio (2009) Large Dataset for Keyphrases Extraction. UNSPECIFIED. (Unpublished)

Download (165Kb) | Preview


    We propose a large dataset for machine learning-based automatic keyphrase extraction. The dataset has a high quality and consist of 2,000 of scientific papers from computer science domain published by ACM. Each paper has its keyphrases assigned by the authors and verified by the reviewers. Different parts of papers, such as title and abstract, are separated, enabling extraction based on a part of an article's text. The content of each paper is converted from PDF to plain text. The pieces of formulae, tables, figures and LaTeX mark up were removed automatically. For removal we have used Maximum Entropy Model-based machine learning and achieved 97.04% precision. Preliminary investigation with help of the state of the art keyphrase extraction system KEA shows keyphrases recognition accuracy improvement for refined texts.

    Item Type: Departmental Technical Report
    Department or Research center: Information Engineering and Computer Science
    Subjects: Q Science > QA Mathematics > QA075 Electronic computers. Computer science
    Uncontrolled Keywords: Keyphrases Extraction, Machine Learning, Large Dataset
    Report Number: DISI-09-055
    Repository staff approval on: 18 Sep 2009

    Actions (login required)

    View Item