Author
Eghbalzadeh, H., Hosseini, B, Khadivi, S., Khodabakhsh, Ali
Publication Date
2012
Publication Place
-
IEEE
Subject
Text classification, Text mining, Categorization, Subject and trend detection
Type
Document
Language
English
Digital
Yes
Manuscript
No
Library
Özyeğin University
Library Asset ID
2-s2.0-84876395015
Record ID
dc04dc2e-213c-4b81-93a6-b706f0428e0e
Date
2012
Sample Text
Lack of multi-application text corpus despite of the surging text data is a serious bottleneck in the text mining and natural language processing especially in Persian language. This paper presents a new corpus for NEWS articles analysis in Persian called Persica. NEWS analysis includes NEWS classification, topic discovery and classification, trend discovery, category classification and many more procedures. Dealing with NEWS has special requirements. First of all it needs a valid and NEWS-content-enriched corpus to perform the experiments. Our Approach is based on a modified category classification and data normalization over Persian NEWS articles which has led to creation of a multipurpose Persian corpus which shows reasonable results in text mining outcomes. In the literature, regarding to our knowledge there are few Persian corpuses but none of them have Persian NEWS time trend characteristics. Empirical results on our benchmark indicate that in addition to reducing the problem dimensions and useless content, Persica keeps admissible validity and reliability in comparison with standard corpuses in the literature.
DOI
10.1109/ISTEL.2012.6483172