Research Article

Benchmarking LLM Tool Orchestration in Resource-Constrained Agentic Workflows: Native MCP Function Calling vs Code Mode

by  Abasiono Mbat, Oselumese Agbonrofo, Oladimeji Abaniwonnda, Samuel Oyefusi
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 136
Published: August 2026
Authors: Abasiono Mbat, Oselumese Agbonrofo, Oladimeji Abaniwonnda, Samuel Oyefusi
10.5120/ijcae10508249e55
PDF

Abasiono Mbat, Oselumese Agbonrofo, Oladimeji Abaniwonnda, Samuel Oyefusi . Benchmarking LLM Tool Orchestration in Resource-Constrained Agentic Workflows: Native MCP Function Calling vs Code Mode. International Journal of Computer Applications. 187, 136 (August 2026), 1-9. DOI=10.5120/ijcae10508249e55

                        @article{ 10.5120/ijcae10508249e55,
                        author  = { Abasiono Mbat,Oselumese Agbonrofo,Oladimeji Abaniwonnda,Samuel Oyefusi },
                        title   = { Benchmarking LLM Tool Orchestration in Resource-Constrained Agentic Workflows: Native MCP Function Calling vs Code Mode },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 136 },
                        pages   = { 1-9 },
                        doi     = { 10.5120/ijcae10508249e55 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Abasiono Mbat
                        %A Oselumese Agbonrofo
                        %A Oladimeji Abaniwonnda
                        %A Samuel Oyefusi
                        %T Benchmarking LLM Tool Orchestration in Resource-Constrained Agentic Workflows: Native MCP Function Calling vs Code Mode%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 136
                        %P 1-9
                        %R 10.5120/ijcae10508249e55
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Large Language Model (LLM) agents are commonly implemented as iterative function-calling loops that enable the use of external tools. The Model Context Protocol (MCP) has emerged as the main example of this paradigm. However, function-calling loops often reduce an agent’s reasoning ability because of increased token usage and high context occupancy. This study compares native MCP function-calling loops with Code Mode under strict engineering constraints. Six scenarios were constructed comparing MCP with a TypeScript Code Mode environment across eight models, including Claude Opus 4.6, GPT-5.4 and Kimi-K2.5. Results show that Code Mode reduced mean token consumption per representative run by 37.6%, reduced mean cost by 40.7% and raised the individual model pass rate from 59.4% to 63.6%, with per-model cost reductions reaching 70.6% for GPT-5.4. However, these gains are not uniform. Weaker models underperform because compiler errors and execution-contract violations trigger retry loops that raise cost, and Gemini-family models regress sharply under Code Mode, passing roughly two thirds of scenarios under native function calling against roughly one third under Code Mode. Once Gemini models are excluded, the Code Mode advantage widens from 4.2 to 16.7%. Overall, Code Mode manages token usage more efficiently and executes tasks successfully when paired with highly capable models, but degrades when paired with less capable ones. These findings suggest that the choice of a tool orchestration paradigm should be an adaptive runtime decision informed by model capability rather than a fixed architectural choice.

References
  • A. Agache, M. Brooker, A. Iordache, F. Liguori, R. Neugebauer, P. Piwonka, and D.-M. Marin. Firecracker: Lightweight virtualization for serverless applications. In Proceedings of the 17th USENIX Conference on File and Storage Technologies, pages 419–434, 2020.
  • Anthropic. Code execution with mcp, 2026. Accessed: February 2, 2026.
  • Anthropic. Model context protocol documentation, 2026. Accessed: January 18, 2026. https://modelcontextprotocol.io/docs/getting-started/intro.
  • L. Chen, M. Zaharia, and J. Zou. Frugalgpt: How to use large language models while reducing cost and improving performance. In Proceedings of the 40th International Conference on Machine Learning (PMLR), volume 202, pages 4822–4846, 2023.
  • W. Chen, X. Chang, T. Schick, and W. Y. Wang. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588, 2022.
  • L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig. Pal: Program-aided language models. In Proceedings of the 40th International Conference on Machine Learning (PMLR), volume 202, pages 10764–10799, 2023.
  • K. Ge, Y. Zhao, J. Sheng, Y. Jiang, C. A. Yang, K. Yang, X. Wang, Y. Wang, et al. Aios: Llm agent operating system. Journal of Machine Learning Research (COLM 2025), 2025.
  • D. Jaroslawicz, B. Whiting, P. Shah, K. Maamari, et al. How many instructions can llms follow at once? arXiv preprint arXiv:2507.11538, 2025.
  • C. E. Jimenez, K. R. Narasimhan, J. Boyreau, R. Yang, Y. Xu, P. Rosand, S. Xia, J. Zhao, et al. Swe-bench: Can language models resolve real-world github issues? In Proceedings of the Twelfth International Conference on Learning Representations, 2024.
  • W. Liu, Y. Zhang, X. Wang, H. Chen, J. Li, Q. Zhao, T. Yang, and Z. Huang. Mcpagentbench: A real-world task benchmark for evaluating llm agent mcp tool use. arXiv preprint arXiv:2512.24565, 2025.
  • X. Liu et al. Agentbench: Evaluating llms as agents. In Proceedings of the Twelfth International Conference on Learning Representations, 2024.
  • OpenAI. Function calling and tool use in the openai api, 2026. Accessed: February 4, 2026. https://platform.openai.com/docs/guides/function-calling.
  • OpenRouter. Unified api for llm providers, 2026. Accessed: February 10, 2026. https://openrouter.ai/docs.
  • C. Packer, V. Fang, S. G. Patil, K. Lin, S. Wooders, and J. E. Gonzalez. Memgpt: Towards llms as operating systems. In Proceedings of the 2024 ACM SIGOPS 29th Symposium on Operating Systems Principles, 2024.
  • S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez. Gorilla: Large language model connected with massive apis. arXiv preprint arXiv:2305.15334, 2023.
  • B. Qiao et al. Taskweaver: A code-first agent framework. arXiv preprint arXiv:2311.17541, 2023.
  • Y. Qin et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. In Proceedings of the Twelfth International Conference on Learning Representations, 2024.
  • T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, volume 36, pages 68539–68551, 2023.
  • K. Varda and S. Pai. Code mode: The better way to use mcp. Cloudflare Blog, 2025. Accessed: January 29, 2026. https://blog.cloudflare.com/code-mode/.
  • F. F. Xu et al. Theagentcompany: Benchmarking llm agents on consequential real world tasks. arXiv preprint arXiv:2412.14161, 2024.
  • S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2023.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Model Context Protocol (MCP) Code Mode LLM Tool Use Agentic Workflows Agent Benchmarking

Powered by PhDFocusTM