Research Article

A Natural Language Processing Approach to Document-based Question Answering

by  Abhijeetsinh Jadeja, Lakdawala Bhumika Jashvantlal, Sandip Patel, Sanket Trivedi
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 138
Published: August 2026
Authors: Abhijeetsinh Jadeja, Lakdawala Bhumika Jashvantlal, Sandip Patel, Sanket Trivedi
10.5120/ijcaa26c135d353e
PDF

Abhijeetsinh Jadeja, Lakdawala Bhumika Jashvantlal, Sandip Patel, Sanket Trivedi . A Natural Language Processing Approach to Document-based Question Answering. International Journal of Computer Applications. 187, 138 (August 2026), 40-43. DOI=10.5120/ijcaa26c135d353e

                        @article{ 10.5120/ijcaa26c135d353e,
                        author  = { Abhijeetsinh Jadeja,Lakdawala Bhumika Jashvantlal,Sandip Patel,Sanket Trivedi },
                        title   = { A Natural Language Processing Approach to Document-based Question Answering },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 138 },
                        pages   = { 40-43 },
                        doi     = { 10.5120/ijcaa26c135d353e },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Abhijeetsinh Jadeja
                        %A Lakdawala Bhumika Jashvantlal
                        %A Sandip Patel
                        %A Sanket Trivedi
                        %T A Natural Language Processing Approach to Document-based Question Answering%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 138
                        %P 40-43
                        %R 10.5120/ijcaa26c135d353e
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

DocuMind is a document-based question answering system built entirely on classical NLP techniques, without relying on deep learning. It combines TF-IDF vectorization with cosine similarity to locate relevant passages within uploaded documents (PDF, TXT, DOCX), then applies rule-based logic to extract precise answers based on question type—WHO, WHEN, WHERE, WHY, HOW MANY, WHAT, and YES/NO. Deployed as a Flask web application, the system delivers real-time answers with confidence scores and links back to source context for verification. Evaluation across diverse documents shows that DocuMind achieves solid retrieval precision and fast memory performance, demonstrating that lightweight, feature-engineered NLP approaches remain a practical and efficient alternative to deep learning for document-based question answering.

References
  • Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
  • Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with Python: analyzing text with the natural language toolkit. " O'Reilly Media, Inc.".
  • Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186).
  • Ferrucci, D., Brown, E., Chu-Carroll, J., Fan, J., Gondek, D., Kalyanpur, A. A., ... & Welty, C. (2010). Building Watson: An overview of the DeepQA project. AI magazine, 31(3), 59-79.
  • Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond (Vol. 4). Now Publishers Inc.
  • Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016, November). Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 conference on empirical methods in natural language processing (pp. 2383-2392).
  • Jones, K. S. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1), 11–21.
  • Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33, 9459-9474.
  • Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  • Manning, C. D. (2008). Introduction to information retrieval. Syngress Publishing.
  • Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., ... & Yih, W. T. (2020, November). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) (pp. 6769-6781).
  • Robertson, S. E., & Spärck Jones, K. (1994). Simple, proven approaches to text retrieval (No. UCAM-CL-TR-356). University of Cambridge, Computer Laboratory.
  • Xu, Y. et al. (2020). LayoutLM: Pre-training of text and layout for document image understanding. KDD 2020.
  • Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  • Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., & Gurevych, I. (2021). Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663.
  • Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
  • Woods, W. A. (1973, June). Progress in natural language understanding: an application to lunar geology. In Proceedings of the June 4-8, 1973, national computer conference and exposition (pp. 441-450).
  • Zhu, F., Lei, W., Wang, C., Zheng, J., Poria, S., & Chua, T. S. (2021). Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv preprint arXiv:2101.00774.
  • Yashoda Suthar, Krina Gondaliya, Umeshkumar Joshi, and Jayashri Patil, trans. 2026. “A Comprehensive Review on Hand Gesture and Facial Expression Analysis for Emotion Recognition”. International Journal of Scientific Research in Science and Technology 13 (4): 35-44. https://doi.org/10.32628/IJSRST26133257.
  • Abhijeetsinh Jadeja, Umeshkumar Joshi, and Chaudhari Vipulkumar Kantilal, “AI Techniques for Indian Legal Document Summarization: A Survey”, Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol, vol. 12, no. 4, pp. 47–55, Jul. 2026, doi: 10.32628/CSEIT261244.
  • Abhijeetsinh Jadeja, Priyanka Ameta, Deepika Ameta, Asha Patil . Depression Severity Classification from Social Media Text using Natural Language Processing and Machine Learning. International Journal of Computer Applications. 187, 116 (Jun 2026), 27-31. DOI=10.5120/ijca9769b912849a.
  • Drashti Bhikadiya, Hemangi Kacha, Abhijeetsinh Jadeja, Jayashri Patil. Cyberbullying Detection using Transformer Architectures: A Comparative Experimental Study. International Journal of Computer Applications. 187, 115 (June 2026), 31-37. DOI=10.5120/ijcae8635085cf6c.
  • Patil, J., Prajapati, K., Patel, D., Chauhan, R., Patel, M. (2026). A Review of Transforming AI for Depression Detection: Transformer Model Dominance, Multimodal Approaches, and Future Pathways. In: Bansal, J.C., Borah, S., Hussain, S., Salhi, S. (eds) Computing and Machine Learning. CML 2025. Lecture Notes in Networks and Systems, vol 1612. Springer, Singapore. https://doi.org/10.1007/978-981-95-2872-1_7.
  • Patil, J., Patil, V., Prajapati, K., Patel, D., Trivedi, S., & Patel, R. (2024, August). Enhanced Depression Detection on Social Media Using Advanced Machine Learning and Linguistic Analysis Techniques. In International Conference on Intelligent.
  • Verma, S., Patil, J. A., & Tamhankar, I. (2024, July). Integrating Natural Language Processing with CGANs to Generate Customized, Realistic Traffic Scenarios for Autonomous Vehicle Training. In 2024 IEEE International Conference on Smart Power Control and Renewable Energy (ICSPCRE) (pp. 1-5). IEEE.
  • Jayashri Patil and Kamini Sharma, “Personality Recognition from Text in Indian Languages : Challenges, Progress, and Future Directions”, Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol, vol. 11, no. 5, pp. 398–401, Oct. 2025, doi: 10.32628/CSEIT251117136.
  • Patil, M. J. A., & Godhwani, M. P. B. (2016). Review of Name Entity Recognition in Marathi Language. IJSART, 2(6).
  • Patil, J., & Sheth, J. (2021). Comparative study of data sources, features, and approaches for automatic personality classification from text. Int. J. Comput. Appl, 174, 0975-8887.
  • Patil, J., & Sheth, J. (2022). Deep Learning and Machine Learning Approaches for the Classification of Personality Traits. In Rising Threats in Expert Applications and Solutions: Proceedings of FICR-TEAS 2022 (pp. 139-146). Singapore: Springer Nature Singapore.
  • Kacha, H., Bhikadiya, D., Sharma, K., & Patil, J. (2025). Comparative Analysis of Visual Motion and Multimodal Strategies in Depression Recognition. International Journal of Computer Applications, 187(58), 58-64.
  • Jayshri Patil, Dr. Jikitsha Sheth Challenges to the Development of Psycholinguistic Dictionary in the Hindi Language Intended for Personality Recognition, IJRTI, Volume 7, Issue 6, 2022, ISSN: 2456-3315.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Answer Extraction Term Frequency–Inverse Document Frequency Angle-Based Matching Data Search Language Processing Tech Web Server Framework Text Insight Systems Segment Finding

Powered by PhDFocusTM