Experimental Data and Evaluation Materials for "An Execution-Aware Evaluation Architecture for AI-Agent Orchestration of Long-Horizon Mobile Manipulation in a Flexible Assembly Environment"
Authors/Creators
Description
This deposit contains the experimental evaluation data and supplementary documentation supporting the manuscript “An Execution-Aware Evaluation Architecture for AI-Agent Orchestration of Long-Horizon Mobile Manipulation in a Flexible Assembly Environment,” submitted to the CIRP Global Web Conference 2026 (CIRPe 2026).
The Excel workbook contains the complete row-level evaluation data for ten multimodal large language models across three scenarios: long-horizon mobile manipulation, vision-grounded attribute and spatial reasoning, and memory-dependent planning. It includes the experimental prompts, expected tool-invocation sequences, evaluation metrics, timing measurements, model-level averages, standard deviations, and manually assigned failure categories.
The deposit also includes machine-readable (tool_schemas.json) and human-readable (tool_schemas.md) documentation of the 24 ROS 2 tools exposed to the AI agents during the experiments. The files describe the registered tool names, functions, parameters, types, default values, required fields, return types, and source locations. Functions and implementation helpers that were not registered with the evaluated agents are excluded.
Two representative session logs from Scenario 1, Prompt 2 are provided. One illustrates successful task completion with corrective behavior. The other documents a cascading unsuccessful execution involving omitted steps, a perception-pipeline failure, continued execution after an unsuccessful grasp, and an incorrect final completion claim.
The two logs are illustrative examples and do not constitute the complete raw-log archive. Complete numerical results and manual error annotations are provided in the Excel workbook. Users should consult the accompanying README for the workbook structure, metric definitions, tool-schema provenance, data-quality notes, and limitations affecting cross-model comparisons.
Files
README.md
Files
(196.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:6d57dcdeeabece12e72e7ef9687787ee
|
3.7 kB | Download |
|
md5:5a6277d446e281fadb158e104bc5a98d
|
3.0 kB | Download |
|
md5:975959c4605c04b30edfb959d2c6d5c7
|
116.6 kB | Download |
|
md5:7605bda4d95efb91d3f495a7331f413d
|
18.0 kB | Preview Download |
|
md5:af9d5db40c5c8fcf1d8feece55fc9a43
|
31.5 kB | Preview Download |
|
md5:7000be582939dcb6b7eb42267e38c056
|
23.8 kB | Preview Download |