There is a newer version of the record available.

Published May 18, 2026 | Version v1

RL-Pretrained Action Experts for Vision-Language-Action Models: Towards Physical Priors in Diffusion Policy Initialization

Authors/Creators

Description

Vision-Language-Action (VLA) models are promising as they allow to interact with a robot through natural language and are very generalist policies. VLAs attach a randomly-initialized Diffusion Transformer (DiT) to a pretrained Vision-Language Model (VLM) and train it via supervised fine-tuning (SFT) on real or synthetic demonstration data. In parallel, the development of modern physics simulator (MuJoCo, Isaac Lab) combined with reinforcement learning (RL) has produced highly capable control policies especially for the locomotion of very complex robots like humanoids or dexterous manipulation. One of the main bottleneck of physical AI is the lack of data. The main advantage of RL training is that it can be run inside a simulator. I propose to use RL to pretrain a VLA in a simulator to let it build a physical knowledge a priori. To do so, it requires to initialize the DiT. I show that this is feasible with two established architectures: a causal transformer trained by standard PPO, or a diffusion policy trained by DPPO or DDiffPG.

Files

paper.pdf

Files (224.1 kB)

Name Size Download all
md5:5c6732fdc6ff373dcd749078a5b0c614
224.1 kB Preview Download