Research Article

Enhancing Patent Readability: Leveraging Large Language Model-Generated Taxonomies for Prior Art Analysis

by  Rohan Kummaraguntla, Andrew J. Ouderkirk
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 131
Published: August 2026
Authors: Rohan Kummaraguntla, Andrew J. Ouderkirk
10.5120/ijca1159a412d939
PDF

Rohan Kummaraguntla, Andrew J. Ouderkirk . Enhancing Patent Readability: Leveraging Large Language Model-Generated Taxonomies for Prior Art Analysis. International Journal of Computer Applications. 187, 131 (August 2026), 1-9. DOI=10.5120/ijca1159a412d939

                        @article{ 10.5120/ijca1159a412d939,
                        author  = { Rohan Kummaraguntla,Andrew J. Ouderkirk },
                        title   = { Enhancing Patent Readability: Leveraging Large Language Model-Generated Taxonomies for Prior Art Analysis },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 131 },
                        pages   = { 1-9 },
                        doi     = { 10.5120/ijca1159a412d939 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Rohan Kummaraguntla
                        %A Andrew J. Ouderkirk
                        %T Enhancing Patent Readability: Leveraging Large Language Model-Generated Taxonomies for Prior Art Analysis%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 131
                        %P 1-9
                        %R 10.5120/ijca1159a412d939
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Patent documents are notoriously difficult to read because of their technical jargon, strict formatting, and lack of semantic structure. This paper studies the application of large language models (LLMs) to produce multi-level hierarchical taxonomies as a strategy to make patents more readable and applicable. By converting unstructured language into structured hierarchies, automated taxonomies offer an intuitive and scalable solution to navigating dense legal text for inventors, researchers, and intellectual property professionals. Patents from various fields including software, medical devices, and materials science were analyzed to evaluate the consistency, depth, and readability of the generated taxonomies. The results indicate that LLMs can effectively restructure complex legal documents into understandable, layered forms— providing an accurate and consistent tool for enhancing information retrieval, prior-art analysis, and a step towards human-artificial intelligence collaboration in intellectual property applications.

References
  • Jin-Dong Li, Kazuaki Seki, and Hisashi Ueda. A novel approach for improving patent readability. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, 2014. URL https://aclanthology. org/W14-1208.pdf.
  • Jiacheng Li, Yi Zhang, Yichi Zhan, Bill Yuchen Lin, and Xiang Ren. Gptpedia: How gpt understands and uses wikipedia. arXiv preprint arXiv:2505.19345, 2024. URL https:// arxiv.org/pdf/2505.19345.
  • Hongyi Fu, Fuwen Luo, Bill Yuchen Lin, Yijia Liu, and Xiang Ren. Are long context language models reading between the lines? arXiv preprint arXiv:2403.04105, 2024. URL https: //arxiv.org/abs/2403.04105.
  • Richard Pak, Scott Pautz, and Robert Iden. Information organization and retrieval: An assessment of taxonomical and tagging systems. Cognitive Technology, 12(1):31–44, 2007.
  • Edward O. Wilson. Taxonomy as a fundamental discipline. Philosophical Transactions of the Royal Society of London B: Biological Sciences, 359(1444):739, 2004. doi: 10.1098/rstb. 2003.1440. URL https://royalsocietypublishing. org/doi/10.1098/rstb.2003.1440.
  • Nishant Garg, Bibek Adhikari, and et al. Semantic embeddings improve the classification and retrieval of patent documents. Scientometrics, 129(2):375–400, 2024. doi: 10.1007/ s11192-024-05206-w. URL https://link.springer. com/article/10.1007/s11192-024-05206-w.
  • Jieh-Sheng Lee and Jieh Hsiang. Patentbert: Patent classification with fine-tuning a pre-trained bert model. arXiv preprint arXiv:1906.02124, 2019. URL https://arxiv.org/abs/ 1906.02124.
  • Shreya Shukla, Nakul Sharma, Manish Gupta, and Anand Mishra. Patentlmm: Large multimodal model for generating descriptions for patent figures. arXiv preprint arXiv:2501.15074, 2025. URL https://arxiv.org/abs/ 2501.15074.
  • Sunyang Fu, David Chen, Huan He, Sijia Liu, Sungrim Moon, Kevin J. Peterson, Feichen Shen, Liwei Wang, YanshanWang, AndrewWen, Yiqing Zhao, Sunghwan Sohn, and Hongfang Liu. Clinical concept extraction: a methodology review, 2019. URL https://arxiv.org/abs/1910.11377.
  • Xi Yang, Jiang Bian, William R. Hogan, et al. Clinical concept extraction using transformers, 2020. URL https: //academic.oup.com/jamia/article-abstract/27/ 12/1935/5943218.
  • Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017.
  • Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020. URL https://arxiv.org/abs/ 2001.03708.
  • Yifan Huang, Yicheng Jiang, Siyuan Wang, and et al. Hierarchical memory transformer for long document understanding. arXiv preprint arXiv:2404.18255, 2024. URL https: //arxiv.org/abs/2404.18255.
  • Asim Abbas, Muhammad Afzal, Jamil Hussain, T. Ali, H.S.M. Bilal, Seokhee Lee, and Seokhee Jeon. Proposed clinical concept extraction methodology, 2021. URL https://www.researchgate.net/figure/ Proposed-clinical-concept-extraction-methodology-process-fig1_355166145.
  • Ziwei Ji et al. Survey of hallucination in natural language generation. ACM Computing Surveys, 2023.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Large Language Models Natural Language Processing Patent Analytics Intellectual Property Information Retrieval Document Readability Taxonomy Generation Prior Art Analysis

Powered by PhDFocusTM