CompToolBench: A Compositional Benchmark for Fine-Grained Tool-Use Evaluation of Large Language Models
Description
We introduce CompToolBench, a benchmark for evaluating LLM tool-use across four composition levels (single, sequential, parallel, and graph) using 200 tasks and 106 real tools drawn from free public APIs. We evaluate 18 models spanning cloud and local deployments and identify a Selection Gap: models that correctly select tools frequently fail to produce valid calls, with an average gap of 13.2 percentage points. Results show that compositional complexity does not uniformly degrade performance, and that local models now approach cloud-model accuracy on structured tool-use tasks.
Files
CompToolBench_Rahman_2026.pdf
Files
(254.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:f268b899126c70b9d25b6778e1ae5087
|
254.8 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/ronyrahmaan/comptoolbench (URL)
Software
- Repository URL
- https://github.com/ronyrahmaan/comptoolbench
- Programming language
- Python