Published March 8, 2026 | Version v1.0

CompToolBench: A Compositional Benchmark for Fine-Grained Tool-Use Evaluation of Large Language Models

Authors/Creators

  • 1. ROR icon Texas Tech University

Description

We introduce CompToolBench, a benchmark for evaluating LLM tool-use across four composition levels (single, sequential, parallel, and graph) using 200 tasks and 106 real tools drawn from free public APIs. We evaluate 18 models spanning cloud and local deployments and identify a Selection Gap: models that correctly select tools frequently fail to produce valid calls, with an average gap of 13.2 percentage points. Results show that compositional complexity does not uniformly degrade performance, and that local models now approach cloud-model accuracy on structured tool-use tasks.

Files

CompToolBench_Rahman_2026.pdf

Files (254.8 kB)

Name Size Download all
md5:f268b899126c70b9d25b6778e1ae5087
254.8 kB Preview Download

Additional details

Related works

Is supplemented by
Software: https://github.com/ronyrahmaan/comptoolbench (URL)

Software

Repository URL
https://github.com/ronyrahmaan/comptoolbench
Programming language
Python