When should a robot keep going—and when should it try a different approach?
A vision-language-action policy can succeed from some configurations and struggle from others. During an execution, continuing the current attempt is only one way to use that policy. A robot can also return to a useful point along its path and start a new continuation.
ACRO makes this decision for a fixed VLA. It estimates success prospects, chooses a promising re-entry state, and physically returns the robot there. The same policy then resumes from a fresh observation.
Across SimplerEnv and RoboCasa, ACRO improves task success under the same physical execution budget, including every return step.
Why orchestrate rollouts?
The outcome of one attempt does not exhaust a policy’s capabilities. An approach can miss a handle, lose a useful contact, or settle into a repetitive motion. A return to a previously visited configuration can give the policy another opportunity to act.
A useful intervention needs three things: a better continuation, a feasible return, and enough remaining time to finish. ACRO brings these decisions into the execution loop.
Actor–Critic orchestration
A Critic learned from uninterrupted policy rollouts estimates the chances of task completion. It helps decide when to intervene and where the policy should resume.
Estimating success
The Q estimate assesses the current state and proposed action chunk. A calibrated rule uses recent Q estimates to activate an intervention. The V estimate ranks candidate states along the executed trajectory and selects a promising re-entry target.
Trajectory Retraction
The robot follows a shortcut inside a tube around positions it has already visited. This removes detours while keeping the return in explored space. Objects retain their current state, and the movement consumes part of the execution budget.
How the Critic is learned
The Critic uses features from the frozen VLA and the robot’s proprioceptive state. A GRU summarizes recent observations. Joint Q and V heads predict success with a piecewise-exponential survival model, so training can use both completed and censored rollouts.
The intervention threshold is calibrated on held-out rollouts. The public code provides the selected Critic, its loss, the Q-history gate, and the recovery loop.
The execution loop
Propose
The fixed VLA proposes an action chunk from the current observation.
Assess
The Critic estimates continuation prospects and monitors recent Q values.
Return
If intervention activates, select a historical state and physically retract.
Resume
Clear the action queue and query the same VLA from the reached observation.
Every policy action and return action counts toward the same physical step limit. The process can repeat while the execution has time remaining.
Results
ACRO is evaluated with GR00T N1.7 and π0.5. Within each setting, the base policy stays fixed and all methods receive the same physical execution budget, including recovery motion.
49.6% → 60.0%
60.7% → 68.0%
GR00T N1.7 · SimplerEnv Bridge
TASK SUCCESS (%)Critic-guided recovery
On RoboCasa, ACRO reaches 68.0% success, compared with 60.7% for the base policy and 59.5% for periodic rewinds. ACRO and periodic rewinds initiate similar numbers of returns, while ACRO spends about half as many physical steps returning.
| Strategy | Success | Returns / trial | Return steps / trial |
|---|---|---|---|
| Base VLA | 60.7% | — | — |
| Periodic rewind | 59.5% | 1.04 | 92.7 |
| ACRO | 68.0% | 1.06 | 46.3 |
Scaling with execution time
Extra execution time gives ACRO opportunities for renewed attempts. Over the final 200 steps of the SimplerEnv budget, ACRO adds 43 completed trials while the base policy adds 25.
More useful attempts from the same policy
ACRO treats rollout execution as a sequence of decisions about continuation and intervention. Learned success estimates and physical re-entry turn the capabilities of a fixed VLA into more completed tasks.
The public package contains the core Critic, intervention gate, retraction planner, and orchestration loop, with a small executable example. The paper provides the evaluation setup and full results.
Cite this work
@misc{acro2026,
title = {{ACRO}: Actor--Critic Rollout Orchestration},
author = {Lee, Junhoo and Kim, Seungyeon and Kim, Baekseung and Kim, Minkyu and Jeon, Suhyun and Kim, Jimyeong and Kwak, Nojun},
year = {2026},
url = {https://junhoo.me/acro}
}