arxiv_ml 92% Match Research Paper Robotics researchers,RL researchers,AI engineers working on embodied agents 1 week ago

$\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

robotics › robotics-rl

📄 Abstract

Abstract: Vision-Language-Action (VLA) models enable robots to understand and perform complex tasks from multimodal input. Although recent work explores using reinforcement learning (RL) to automate the laborious data collection process in scaling supervised fine-tuning (SFT), applying large-scale RL to flow-based VLAs (e.g., $\pi_0$, $\pi_{0.5}$) remains challenging due to intractable action log-likelihoods from iterative denoising. We address this challenge with $\pi_{\text{RL}}$, an open-source framework for training flow-based VLAs in parallel simulation. $\pi_{\text{RL}}$ implements two RL algorithms: (1) {Flow-Noise} models the denoising process as a discrete-time MDP with a learnable noise network for exact log-likelihood computation. (2) {Flow-SDE} integrates denoising with agent-environment interaction, formulating a two-layer MDP that employs ODE-to-SDE conversion for efficient RL exploration. We evaluate $\pi_{\text{RL}}$ on LIBERO and ManiSkill benchmarks. On LIBERO, $\pi_{\text{RL}}$ boosts few-shot SFT models $\pi_0$ and $\pi_{0.5}$ from 57.6% to 97.6% and from 77.1% to 98.3%, respectively. In ManiSkill, we train $\pi_{\text{RL}}$ in 320 parallel environments, improving $\pi_0$ from 41.6% to 85.7% and $\pi_{0.5}$ from 40.0% to 84.8% across 4352 pick-and-place tasks, demonstrating scalable multitask RL under heterogeneous simulation. Overall, $\pi_{\text{RL}}$ achieves significant performance gains and stronger generalization over SFT-models, validating the effectiveness of online RL for flow-based VLAs.

Authors (13)

Kang Chen

Zhihao Liu

Tonghe Zhang

Zhen Guo

Si Xu

Hao Lin

+7 more

Submitted

October 29, 2025

arXiv Category

cs.LG

arXiv PDF Code

Key Contributions

Introduces $\pi_{\text{RL}}$, an open-source framework for training flow-based VLAs using RL in parallel simulation. It proposes two novel RL algorithms: Flow-Noise (discrete-time MDP) and Flow-SDE (two-layer MDP with ODE-to-SDE conversion), which address the challenge of intractable action log-likelihoods and enable efficient RL exploration for these models.

Business Value

Enables more efficient and effective training of robots capable of understanding and executing complex tasks based on visual and language instructions, accelerating the development of autonomous systems.

Paper Metadata

Innovation Type

Algorithmic Framework

Deployment Feasibility

The framework facilitates training, but deployment requires robust robotic hardware and sim-to-real transfer capabilities.

Limitations Addressed

Intractable action log-likelihoods in flow-based VLAs for RL,Difficulty in applying large-scale RL to these models,High cost of supervised data collection

View Code on GitHub

Technical Tags

Vision-Language-Action (VLA)Reinforcement Learning (RL)Flow-based ModelsDenoising Diffusion Probabilistic ModelsRobotics ControlSimulationMDPSDE

Research Topics

RL for robotic controlVision-language grounding for actionEfficient RL training for flow modelsSim-to-real transfer

Methods & Architectures

$\pi_{\text{RL}}$ frameworkFlow-Noise RL algorithmFlow-SDE RL algorithmDiscrete-time MDP formulationTwo-layer MDP formulationODE-to-SDE conversion Flow-based Vision-Language-Action (VLA) modelsDenoising Diffusion Probabilistic Models

Applications & Tasks

Robotics Embodied AI Human-Robot Interaction Intractable action log-likelihoods in flow-based VLAs for RLChallenges in applying large-scale RL to flow-based VLAsLaborious data collection for supervised fine-tuning Training flow-based VLAs using RLAutomating robotic task executionEnabling robots to perform complex tasks from multimodal input

Datasets & Benchmarks

Datasets

LIBERO, ManiSkill

Related Fields

RoboticsReinforcement LearningComputer VisionNatural Language ProcessingGenerative ModelsSimulation

Keywords

Vision-Language-ActionReinforcement LearningFlow-based ModelsDiffusion ModelsRoboticsSimulationMDPSDEEmbodied AITask ExecutionRobotic ControlOpen Source

Academic Context

#RL for robotic control#Vision-language grounding for action#Efficient RL training for flow models#Sim-to-real transfer

Companies & Organizations

Companies Mentioned

Google

Technology Stack

Frameworks & Libraries

PyTorch

Programming Languages

Python

Commercial Potential

Potential Products

Robotic control softwareAI assistants for industrial automationAutonomous navigation systems

Target Industries

ManufacturingLogisticsWarehousingAutomotiveConsumer Electronics

Use Case Examples

Robots performing assembly tasks based on instructionsAutonomous delivery robotsAssistive robots in homes or workplaces

Competitive Edge

Provides a novel RL framework specifically designed for flow-based VLAs, overcoming key technical hurdles that limit the application of RL in this domain.

Market Opportunity

Rapidly growing market for robotics and autonomous systems.

Revenue Models

Licensing of the frameworkdevelopment of specialized robotic solutionsAI-as-a-service for robot control.

Resource Requirements

Compute Needs

High for training in simulation.

Data Requirements

Simulation environments and potentially real-world robot interaction data.

Deployment Constraints

Requires robust sim-to-real transfer, hardware integration, and safety protocols for robotic deployment.

Scalability

The framework is designed for large-scale RL training in simulation.

Regulatory Considerations

Safety standards for roboticsAI ethics in autonomous systems

Production Readiness

Maturity Level

Research

Time to Market

3-5 years for robust industrial applications.

Licensing

Open Source (Apache 2.0)

Patent Potential

Moderate, for the novel RL algorithms and framework.

View Full Paper Back to Papers