Large Dataset for Keyphrases Extraction

Krapivin, Mikalai and Autaeu, Aliaksandr and Marchese, Maurizio (2009) Large Dataset for Keyphrases Extraction. UNSPECIFIED. (Unpublished)

Preview

PDF
Download (165Kb) | Preview

Abstract

We propose a large dataset for machine learning-based automatic keyphrase extraction. The dataset has a high quality and consist of 2,000 of scientific papers from computer science domain published by ACM. Each paper has its keyphrases assigned by the authors and verified by the reviewers. Different parts of papers, such as title and abstract, are separated, enabling extraction based on a part of an article's text. The content of each paper is converted from PDF to plain text. The pieces of formulae, tables, figures and LaTeX mark up were removed automatically. For removal we have used Maximum Entropy Model-based machine learning and achieved 97.04% precision. Preliminary investigation with help of the state of the art keyphrase extraction system KEA shows keyphrases recognition accuracy improvement for refined texts.

Item Type:	Departmental Technical Report
Department or Research center:	Information Engineering and Computer Science
Subjects:	Q Science > QA Mathematics > QA075 Electronic computers. Computer science
Uncontrolled Keywords:	Keyphrases Extraction, Machine Learning, Large Dataset
Report Number:	DISI-09-055
Repository staff approval on:	18 Sep 2009

Actions (login required)

View Item

Università degli Studi di Trento

Unitn-eprints.PhD

Large Dataset for Keyphrases Extraction

Abstract

Actions (login required)