Research Article

Open Source Autonomous Bengali Corpus

by  Summit Haque, Md. Abu Shahriar Ratul, Md. Yousuf Ali Khan
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 176 - Issue 17
Published: Apr 2020
Authors: Summit Haque, Md. Abu Shahriar Ratul, Md. Yousuf Ali Khan
10.5120/ijca2020920120
PDF

Summit Haque, Md. Abu Shahriar Ratul, Md. Yousuf Ali Khan . Open Source Autonomous Bengali Corpus. International Journal of Computer Applications. 176, 17 (Apr 2020), 33-37. DOI=10.5120/ijca2020920120

                        @article{ 10.5120/ijca2020920120,
                        author  = { Summit Haque,Md. Abu Shahriar Ratul,Md. Yousuf Ali Khan },
                        title   = { Open Source Autonomous Bengali Corpus },
                        journal = { International Journal of Computer Applications },
                        year    = { 2020 },
                        volume  = { 176 },
                        number  = { 17 },
                        pages   = { 33-37 },
                        doi     = { 10.5120/ijca2020920120 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2020
                        %A Summit Haque
                        %A Md. Abu Shahriar Ratul
                        %A Md. Yousuf Ali Khan
                        %T Open Source Autonomous Bengali Corpus%T 
                        %J International Journal of Computer Applications
                        %V 176
                        %N 17
                        %P 33-37
                        %R 10.5120/ijca2020920120
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Through Sentiment Analysis System it is possible to know what kind of information is there in a text. For example, one can identify is the text about a particular product, political view, sport, entertainment, education, politics, etc. or not. It is also possible to further categorize text in positive, negative or neutral. So, through proper Sentiment Analysis, the current technology would go to another step. There are so many works on Sentiment Analysis that have been done already in different languages. But due to lack of data, the work on Sentiment Analysis on Bangla Text is very limited. Because word categorization accuracy depends heavily on the size of the text corpus used to derive the inter-word statistics. So, it was planned to develop an automated corpus generation system that traverses the Web collecting text and stores them under the defined category. This flexible scheme can produce very large general-purpose corpora or particular samples of domain-specific text.

References
  • Sebastiani, F. Text Categorization.
  • Xiao, R. Corpus Creation.
  • Majumder, K. M. Y. A., Islam, M. Z., Zaman, N. U., and Khan, M. Analysis of and Observations from a Bangla News Corpus.
  • Mumin, M. A. A., Shoeb, A. A. M., Selim, M. R., and Iqbal, M. Z. SUMono: A Representative Modern Bengali Corpus
  • McEnery, A., Xiao, R., Tono, Y. Corpora Survey.
  • Panunzi, A., Fabbri, M, Moneglia, M., Gregori, L., and Paladini, S. RIDIRE-CPI: an Open Source Crawling and Processing Infrastructure for web Corpora Building.
  • Sarkar, A. I., Pavel, D. S. H., and Khan, M. Automatic Bangla Corpus Creation.
  • Lesher, G. W. A Web-Based System for Autonomous Text Corpus Generation.
  • Oliver, A. Automaticcreation of WordNets from parallel corpora
  • Pavel Kr´al, P., and Cerisara, C. Automatic Dialog Act Corpuscreation From Web Pages
  • Jha, M., Andreas, J., Thadani, K., Rosenthal, S., and McKeown, K. Corpus Creation for New Genres: A Crowdsourced Approach to PP Attachment
  • Maeda, K., Lee, H., Medero, S., Medero, J., Parker, R., and Strassel, S. Annotation Tool Development for Large-Scale Corpus Creation Projects at the Linguistic Data Consortium.
  • Cieri, C., and Liberman, M. Issues in Corpus Creation and Distribution: The Evolution of the Linguistic Data Consortium
  • Pavel, D. S. H., Sarkar, A. I., and Khan, M. A Proposed Automated Extraction Procedure Of Bangla Text For Corpus Creation In Unicode
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Corpus Autonomous Corpus Autonomous Bengali Corpus ZIPF Law.

Powered by PhDFocusTM