Dominic Rigby

RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Date: 28th September 2026

arXiv Link PDF

Related: π0.5 VLA

Key Points

Key Methods

  1. Train a VLA to output a rich RL-token.
    • This is created by training an small transformer autoencoder on the VLAs embeddings on task relevant data.
    • The aim is to produce a vector which contains all the context the VLA has absorbed
    • Tuned VLA is then frozen
  2. Train small network actor-critic with RL Token and outputted action as observation:
    • RL-token contains compact reasoning info.
    • Outputted action means that we can tune the already generated action, rather than having to learn action dynamics from scratch
      • Action is randomly masked out so stopped the policy just regurgitating the action.
    • Operates over action chunks
    • Policy and critic operate over smaller action chunks, to be more reactive
    • Trained with TD3
    • Light weight MLP actor and critic (256x256 MLPs)