Research Article

AI-Driven Digitization of Marathi Indian Knowledge Systems: A Conceptual Framework for Sustainable and Holistic Development

by  Ujwala Mahajan, Mohit Mahajan
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 134
Published: August 2026
Authors: Ujwala Mahajan, Mohit Mahajan
10.5120/ijca25685698fffb
PDF

Ujwala Mahajan, Mohit Mahajan . AI-Driven Digitization of Marathi Indian Knowledge Systems: A Conceptual Framework for Sustainable and Holistic Development. International Journal of Computer Applications. 187, 134 (August 2026), 42-48. DOI=10.5120/ijca25685698fffb

                        @article{ 10.5120/ijca25685698fffb,
                        author  = { Ujwala Mahajan,Mohit Mahajan },
                        title   = { AI-Driven Digitization of Marathi Indian Knowledge Systems: A Conceptual Framework for Sustainable and Holistic Development },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 134 },
                        pages   = { 42-48 },
                        doi     = { 10.5120/ijca25685698fffb },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Ujwala Mahajan
                        %A Mohit Mahajan
                        %T AI-Driven Digitization of Marathi Indian Knowledge Systems: A Conceptual Framework for Sustainable and Holistic Development%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 134
                        %P 42-48
                        %R 10.5120/ijca25685698fffb
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Indian Knowledge Systems (IKS), which integrate philosophy, science, ecology, and holistic well-being, comprise an extensive collection of traditional wisdom developed over centuries in the Indian subcontinent. A significant portion of this knowledge is preserved in unstructured textual formats, such as manuscripts, traditional literature, and oral documentation, and is found in regional languages like Marathi. Much of this knowledge is still unavailable in the digital age due to a lack of systematic digitization and computational processing. Computational frameworks that can organize, extract, and interpret original knowledge are therefore becoming increasingly necessary. This paper presents a conceptual AI-driven framework for digitizing and organizing Indian Knowledge Systems found in Marathi. The framework includes text preprocessing, morphological analysis, Named Entity Recognition (NER), Multiword Expression (MWE) detection, semantic concept extraction, and thematic clustering. These NLP techniques enable automatic identification and arrangement of knowledge about sustainability, ethics, ecology, and community-based development. The suggested model places a strong emphasis on combining modern AI technologies with knowledge structures that originate in culture. The framework offers a scalable method for creating digital archives of traditional knowledge by permitting semantic interpretation and knowledge representation of Marathi texts. Initiatives for sustainable development, education, research, and cultural preservation can all benefit from such repositories. By offering a conceptual architecture for maintaining and restoring traditional knowledge systems through computational techniques, the study adds to the growing integration of artificial intelligence and original knowledge digitization.

References
  • Aggarwal, C. C., & Zhai, C. (2012). Mining text data. Springer.
  • Bender, E. M., & Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6, 587–604.
  • Bhattacharyya, P. (2015). IndoWordNet: A lexical resource for Indian languages. In Proceedings of the Global WordNet Conference.
  • Bhattacharyya, P. (2017). Natural language processing for Indian languages. Springer.
  • Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with Python. O’Reilly Media.
  • Chaudhari, S., Patil, S., & Joshi, A. (2023). Hybrid approaches for named entity recognition in low-resource languages.
  • Natural Language Engineering, 29(2), 325–341.
  • Chen, D., & Manning, C. D. (2014). A fast and accurate dependency parser using neural networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  • Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT.
  • Goldberg, Y. (2017). Neural network methods for natural language processing. Morgan & Claypool.
  • Gruber, T. R. (1995). Toward principles for the design of ontologies used for knowledge sharing. International Journal of Human-Computer Studies, 43(5–6), 907–928.
  • Joshi, A., Patil, S., Kale, P., & Bhattacharyya, P. (2022). L3Cube-MahaCorpus and MahaBERT: Marathi monolingual corpus and language model. In Proceedings of the Language Resources and Evaluation Conference (LREC).
  • Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the Association for Computational Linguistics (ACL).
  • Jurafsky, D., & Martin, J. H. (2020). Speech and language processing (3rd ed.). Stanford University.
  • Kumar, A., Singh, R., & Sharma, P. (2021). Cross-lingual natural language processing models for low-resource languages.
  • Artificial Intelligence Review, 54(6), 4125–4145.
  • Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K., & Dyer, C. (2016). Neural architectures for named entity recognition. In Proceedings of NAACL-HLT.
  • Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  • Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
  • Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
  • Navigli, R., & Ponzetto, S. P. (2012). BabelNet: The automatic construction, evaluation, and application of a wide-coverage multilingual semantic network. Artificial Intelligence, 193, 217–250.
  • Patil, S., Joshi, A., Kale, P., & Bhattacharyya, P. (2022). L3Cube-MahaNER: Named entity recognition dataset and baseline models for Marathi. In Proceedings of the Language Resources and Evaluation Conference (LREC).
  • Pennington, J., Socher, R., & Manning, C. (2014). GloVe: Global vectors for word representation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  • Ruder, S. (2019). Neural transfer learning for natural language processing. PhD Thesis, National University of Ireland.
  • Sarkar, D. (2016). Text analytics with Python: A practical, real-world approach to gaining actionable insights from your data. Apress.
  • Saurav, S., Singh, A., & Bhattacharyya, P. (2020). Word embeddings for Indian languages: A survey. ACM Computing Surveys, 53(5), 1–38.
  • Singh, A., & Kumar, R. (2021). Artificial intelligence in cultural heritage preservation: Opportunities and challenges. Journal of Cultural Heritage Management and Sustainable Development, 11(3), 412–428.
  • Smith, A. (2011). Preservation in the digital age. Library Trends, 59(3), 333–352.
  • UNESCO. (2017). Safeguarding intangible cultural heritage in the digital era. UNESCO Publishing.
  • Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS).
  • Zhang, Y., Jin, R., & Zhou, Z. H. (2010). Understanding the bag-of-words model: A statistical framework. International Journal of Machine Learning and Cybernetics, 1(1–4), 43–52.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Indian Knowledge Systems Artificial Intelligence Marathi NLP Multiword Expression Named Entity Recognition Knowledge Digitization

Powered by PhDFocusTM