An Efficient Approach to Machine Learning Based Text Classification Through Distributed Computing
Author | : Raghu Nandan Immaneni |
Publisher | : |
Total Pages | : 75 |
Release | : 2015 |
ISBN-10 | : 1339214954 |
ISBN-13 | : 9781339214955 |
Rating | : 4/5 (54 Downloads) |
Book excerpt: Abstract: Text classification is one of the classical problems in computer science, which is primarily used for categorizing data, spam detection, anonymization, information extraction, text summarization etc. Given the large amounts of data involved in the above applications, automated and accurate training models and approaches to classify data efficiently are needed. In this thesis, an extensive study of the interaction between natural language processing, information retrieval and text classification has been performed. A case study named "keyword extraction" that deals with 'identifying keywords and tags from millions of text questions' is used as a reference. Different classifiers are implemented using MapReduce paradigm on the case study and the experimental results are recorded using two newly built distributed computing Hadoop clusters. The main aim is to enhance the prediction accuracy, to examine the role of text pre-processing for noise elimination and to reduce the computation time and resource utilization on the clusters.