ACROOCTOBER 2026

ROBOT LEARNING · ROLLOUT ORCHESTRATION

ACRO: Actor–Critic
Rollout Orchestration

Give a fixed VLA a more promising continuation.

Junhoo Lee2, Seungyeon Kim1, Baekseung Kim1, Minkyu Kim1, Suhyun Jeon1, Jimyeong Kim1, Nojun Kwak1
1 Seoul National University   2 KAIST
A RECOVERY, STEP BY STEPOPEN STAND MIXER HEAD
01 / RECORDED EXECUTION
20 fps playback
02 / ESTIMATE AND INTERVENE
Continue0 / 224 steps

The frozen policy approaches the mixer. ACRO estimates the prospects of continuing with its proposed actions.

10 0130180224 Q · continuation success
Policy executionPhysical return
03 / RETURN AND RESUME
High-resolution replay frame at the configuration where VLA execution resumes

A fresh observation.
The same VLA.

The robot returns along its visited path, then makes a new approach and opens the mixer head.

Return motion
50 steps
After resuming
44 steps to success

An ACRO execution from the qualitative example in the paper. The video shows the recorded camera view; the right panel shows a high-resolution replay of the same actions. Return motion is included in the physical step count.

When should a robot keep going—and when should it try a different approach?

A vision-language-action policy can succeed from some configurations and struggle from others. During an execution, continuing the current attempt is only one way to use that policy. A robot can also return to a useful point along its path and start a new continuation.

ACRO makes this decision for a fixed VLA. It estimates success prospects, chooses a promising re-entry state, and physically returns the robot there. The same policy then resumes from a fresh observation.

Across SimplerEnv and RoboCasa, ACRO improves task success under the same physical execution budget, including every return step.

1.

Why orchestrate rollouts?

The outcome of one attempt does not exhaust a policy’s capabilities. An approach can miss a handle, lose a useful contact, or settle into a repetitive motion. A return to a previously visited configuration can give the policy another opportunity to act.

Stand mixer recovery sequence: declining continuation estimates, physical retraction, renewed approach and success
ACRO intervenes as continuation prospects decline, returns to an earlier configuration, and resumes the fixed policy. The mixer head opens after the new approach.

A useful intervention needs three things: a better continuation, a feasible return, and enough remaining time to finish. ACRO brings these decisions into the execution loop.

2.

Actor–Critic orchestration

A Critic learned from uninterrupted policy rollouts estimates the chances of task completion. It helps decide when to intervene and where the policy should resume.

ACRO method: assess a proposed action with Q, evaluate historical re-entry states with V, retract along the visited path, and resume the VLA
The Critic evaluates continuation and re-entry prospects. The Actor carries out the return and gives the VLA a new starting observation.

Estimating success

The Q estimate assesses the current state and proposed action chunk. A calibrated rule uses recent Q estimates to activate an intervention. The V estimate ranks candidate states along the executed trajectory and selects a promising re-entry target.

Trajectory Retraction

The robot follows a shortcut inside a tube around positions it has already visited. This removes detours while keeping the return in explored space. Objects retain their current state, and the movement consumes part of the execution budget.

How the Critic is learned

The Critic uses features from the frozen VLA and the robot’s proprioceptive state. A GRU summarizes recent observations. Joint Q and V heads predict success with a piecewise-exponential survival model, so training can use both completed and censored rollouts.

The intervention threshold is calibrated on held-out rollouts. The public code provides the selected Critic, its loss, the Q-history gate, and the recovery loop.

3.

The execution loop

01

Propose

The fixed VLA proposes an action chunk from the current observation.

02

Assess

The Critic estimates continuation prospects and monitors recent Q values.

03

Return

If intervention activates, select a historical state and physically retract.

04

Resume

Clear the action queue and query the same VLA from the reached observation.

Every policy action and return action counts toward the same physical step limit. The process can repeat while the execution has time remaining.

4.

Results

ACRO is evaluated with GR00T N1.7 and π0.5. Within each setting, the base policy stays fixed and all methods receive the same physical execution budget, including recovery motion.

SIMPLERENV+10.4 pp

49.6% → 60.0%

ROBOCASA+7.3 pp

60.7% → 68.0%

GR00T N1.7 · SimplerEnv Bridge

TASK SUCCESS (%)
Task success rates

Critic-guided recovery

On RoboCasa, ACRO reaches 68.0% success, compared with 60.7% for the base policy and 59.5% for periodic rewinds. ACRO and periodic rewinds initiate similar numbers of returns, while ACRO spends about half as many physical steps returning.

Recovery strategy comparison on RoboCasa
StrategySuccessReturns / trialReturn steps / trial
Base VLA60.7%——
Periodic rewind59.5%1.0492.7
ACRO68.0%1.0646.3

Scaling with execution time

Extra execution time gives ACRO opportunities for renewed attempts. Over the final 200 steps of the SimplerEnv budget, ACRO adds 43 completed trials while the base policy adds 25.

Task completion grows with physical step budget for ACRO and Base VLA on SimplerEnv
Success as a function of physical execution steps. Every retraction step is included.
5.

More useful attempts from the same policy

ACRO treats rollout execution as a sequence of decisions about continuation and intervention. Learned success estimates and physical re-entry turn the capabilities of a fixed VLA into more completed tasks.

The public package contains the core Critic, intervention gate, retraction planner, and orchestration loop, with a small executable example. The paper provides the evaluation setup and full results.

Cite this work

BIBTEX
@misc{acro2026,
  title = {{ACRO}: Actor--Critic Rollout Orchestration},
  author = {Lee, Junhoo and Kim, Seungyeon and Kim, Baekseung and Kim, Minkyu and Jeon, Suhyun and Kim, Jimyeong and Kwak, Nojun},
  year = {2026},
  url = {https://junhoo.me/acro}
}