There is a newer version of the record available.

Published July 22, 2026 | Version v1

Data for Arnold: A multi-task, multi-embodiment muscle transformer policy

  • 1. EPFL
  • 2. ROR icon École Polytechnique Fédérale de Lausanne

Description

Supplementary Data — Trained Model Checkpoints and Benchmark Results
====================================================================

This archive set accompanies an anonymous manuscript submission. It contains
the final trained model weights and the corresponding benchmark evaluation
results for every model reported in the paper.

Contents
--------
1. arnold-final-checkpoints.tar.gz
     Trained model weights. Unpacks to:  final_checkpoints/
     ~2.8 GB. One subdirectory per model (see "Models" below), each
     containing per-seed subdirectories (seed_0, seed_1, ...).

     A typical seed directory contains:
       - rl_model_<steps>_steps.zip          Stable-Baselines3 policy checkpoint
       - rl_model_vecnormalize_<steps>_steps.pkl   VecNormalize statistics
       - model_config.json / args.json       training + model configuration
       - <Task>_config.json                  per-task environment configs
       - vocabulary.json                     sensorimotor vocabulary (where used)
       - main_bc_ppo_multi_task.py           training entry point (snapshot)
       - PPO_0/                              training logs / tensorboard events

2. arnold-benchmark-results.tar.gz
     Benchmark evaluation results. Unpacks to:  benchmarks/final_checkpoints/
     ~4 MB. Same model/seed layout as above; each seed directory holds:
       - seed_<n>_<checkpoint>_results.json  per-task benchmark scores
       - *.log                               evaluation run logs

Model

Description

arnold

Proposed method (multi-task compositional policy)

obc

On-policy behavior cloning baseline

obc_10_tasks

OBC trained on the 10-task subset

obc_m / obc_s / obc_xs

OBC capacity variants (medium / small / extra-small)

obc_st

OBC single-task specialists

obc_task_sv

OBC, task-specific sensorimotor vocabulary variant

obc_ppo

OBC with PPO fine-tuning

obc_wo_obs_norm

OBC ablation without observation normalization

bc

Behavior-cloning baseline

ppo_t

Transformer-based PPO (no sensorimotor vocabulary)

ppo_t_sv

Transformer-based PPO with sensorimotor vocabulary

mt-ppo

Multi-task PPO

mt-sac

Multi-task SAC

Files

ZENODO_README.txt

Files (2.8 GB)

Name Size
md5:63fd5efaaa40500408d5821aecd6e1c8
382.0 kB Download
md5:6c5fa8da01f47ad0d5a17ce21a8b5bdb
2.8 GB Download
md5:db00505dfe8debee8053a84769ed491b
2.0 kB Preview Download

Additional details

Related works

Is supplement to
Preprint: arXiv:2508.18066 (arXiv)

Funding

Swiss National Science Foundation
310030_212516
Simons Foundation
SFI-AN-NC-SCN-00007276-14