Published March 16, 2020 | Version v1

Robust Adversarial Imitation Learning from Observation

Description

Imitation learning has been successfully used in many real-world scenarios such as autonomous driving, especially learning from video demonstrations. Unlike traditional imitation learning where both the observations and actions from an expert agent are provided to the student agent for training, imitation learning from observation (ILO) learns from observation only, which is more applicable as it is more intuitive and the actions are usually hard to detect accurately. However, one may encounter gaps in performance between the simulation results and the actual real-world results due to the discrepancy between training and test environments. In this paper, we propose a novel framework named Robust Imitation Learning from Observation (RILO) that aims to provide robustness for the student agent in an ILO setting. We introduce an adversary agent that competes with the student agent and maximally destabilizes the student agent by applying disturbances to the system. We jointly train the agent and the adversary simultaneously so that the student agent is reinforced and becomes more robust to the various conditions. We empirically demonstrate that RILO enhances the robustness of the system, in terms of its steady performance in test environments with different mass values from the training environment.

Files

Files (576.3 kB)

Name Size Download all
md5:fde3d116328accf2951dbdb4ce7adf7f
576.3 kB Download