CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning

Anonymous Authors
Paper under double-blind review at ICLR 2027

CAPEX turns a frozen multimodal foundation model into an autonomous robot demonstrator, and learns from its own execution experience when it actually needs to reason again.

4.3×
more successful demonstrations
12.5% → 53.6% on 360 RoboCasa starts
−80%
API cost per demonstration
$9.15 → $1.83
244 → 37
model requests per demonstration
39.7 → 6.6 min of collection
45% → 95%
collection success on the bimanual YAM
real robot, 2 tasks

Abstract

Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators.

As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data.

We find that CAPEX increases the number of successful demonstrations by 4.3× while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience.

Method

Querying a frontier model after every small motion couples data collection directly to model latency and inference cost. CAPEX (Calibrated Adaptive Planning through EXperience Collection) instead lets the model propose temporally extended plans, and uses the robot's own execution outcomes to learn how much of each plan can be trusted before reasoning again. The foundation model stays frozen throughout; all adaptation happens through a small, bounded cross-episode state.

CAPEX overview diagram comparing human teleoperation, a base VLM that queries once per waypoint, and CAPEX's calibrated multi-waypoint plans with cross-episode experience.
Overview. (a) Teleoperation yields human demonstrations. (b) The base arm queries a frozen VLM once per waypoint. (c) CAPEX's VLM proposes a multi-waypoint plan with stated confidences qk; the robot executes it until a later waypoint's calibrated probability p̂k falls below τ or execution stops, then queries the VLM again. (d) Before each episode, all past attempts are summarized into a cross-episode state: calibrator weights that set how much of a plan runs, and a VLM context holding the robot's own successes and per-phase reliability.
1

Waypoint plans

Each model decision proposes a variable-length plan of up to 24 end-effector waypoints, each with a stated probability of being reached. A locate_pixel tool grounds image pixels to 3-D points.

2

Calibrated commitment

Every executed waypoint is labeled reached or stalled / contact-stopped by the controller, which needs no human annotation. A logistic calibrator refit before each episode decides how long a prefix of the plan to execute before observing again.

3

Learning from experience

The VLM context carries a phase-level reliability summary and up to two of the robot's own previous successes, with actions expressed relative to grounded scene points. The context stays bounded no matter how much is collected.

Real-World Results

Bimanual YAM: autonomous demonstration collection

On a bimanual I2RT YAM, CAPEX succeeds on 40/42 collection attempts versus 18/40 for the base GPT-6 Astra teacher, using 11.7 instead of 37.6 model calls and $0.65 instead of $1.64 per successful demonstration. Both teachers end every attempt by declaring the task done; the operator rejected 22 of the base teacher's claims and 2 of CAPEX's. Each clip shows the front camera, with the left and right wrist cameras stacked on its right.

Bread + corn to plate (same start state)

Places both items without moving the plate.

Pushes the plate off its position, then reports the task done.

In-air bowl handover (same start state)

Lift by the rim, hand over in the air, set down upright on the right.

Same start state; the operator labelled this attempt a failure. Base: 6/20 successes on this task.

Franka Panda: policies trained only on CAPEX data

We train Diffusion Policy and ACT on 20 CAPEX demonstrations per task and compare against the same policies trained on 20 matched human-teleoperated demonstrations. CAPEX-trained policies match or exceed their human-trained counterparts in five of six task–policy comparisons. The rollouts below are closed-loop Diffusion Policy executions with no foundation model in the loop.

Pick Coke “grab the coke can and lift it up” · 4/10 trials succeeded

Stack Cups “stack the orange cup onto the green cup” · 2/10 trials succeeded

Stack Blocks “stack the blue block on top of the green block” · 2/10 trials succeeded

The demonstrations those policies were trained on (collected by CAPEX; exterior and wrist cameras)

Pick Coke: 20/35 attempts succeeded

Stack Cups: 20/20 attempts succeeded

Stack Blocks: 22/28 attempts succeeded (20 kept)

Physical-robot evaluation. Left: autonomous collection on the bimanual YAM. Right: downstream policy success on the Franka (successes / 10 trials).
YAM taskSuccess$ / success ↓ Franka taskDiffusion PolicyACT
BaseCAPEXBaseCAPEX OursHumanOursHuman
Bread + Corn12/2020/211.100.43 Pick Coke4/105/103/103/10
Bowl Handover6/2020/212.730.88 Stack Cups2/100/101/100/10
Stack Blocks2/101/100/100/10
Overall18/4040/421.640.65 Overall8/306/304/303/30

Simulation: RoboCasa

Autonomous demonstration collection

We recreate the initial states of the first 20 official human demonstrations for each of 18 atomic RoboCasa tasks (360 paired starts) and compare a base VLM that predicts one waypoint per request with the same VLM under CAPEX. With GPT-6 Astra, CAPEX outperforms the base teacher on 14 of 18 tasks, 11 of them significantly.

Bar chart of collection success: CAPEX outperforms Base across all five RoboCasa task families; 193/360 vs 45/360 overall with GPT-6 Astra.
Collection success on the same 20 starts per task. Left: GPT-6 Astra by task family. Right: all 18 tasks for both backbones. Whiskers are 95% Wilson intervals.
Every attempt counts toward cost: totals over all attempts, divided by the number of successful demonstrations. Qwen runs locally (no API price).
MethodSuccess ↑$ / success ↓Requests / success ↓Minutes / success ↓
Astra (Base)45/360  12.5%9.1524439.7
Astra + CAPEX193/360  53.6%1.83376.6
Qwen3.8-27B (Base)0/360  0.0%—∞∞
Qwen3.8-27B + CAPEX5/350  1.4%—4,8192,332

Demonstrations collected by CAPEX vs. human teleoperation from the same start

Both sources are rendered by the same RoboCasa pipeline at the same resolution; all simulation clips play in real time.

“Pick the mango from the plate and place it in the pan.”

“Close the toaster oven door.”

“Press the start button on the microwave.”

More CAPEX demonstrations across task families

Close Fridge

Open Cabinet

PnP Counter→Cabinet

Turn Off Stove

PnP Sink→Counter

Navigate Kitchen

Learning from experience

Clearing memory before each episode isolates the effect of experience. Once the robot has at least one prior success, collection success rises from 54.0% to 65.0% on the same starts (p = 0.0016), and successful executions become shorter: −27% model requests and −21% wall-clock time. Raw VLM confidence is poorly calibrated (it never goes below 0.5, while only 70% of waypoints are reached). Calibrating on the robot's own outcomes brings ECE from 0.24 to 0.05 and AUROC from 0.57 to 0.75, so the robot can safely truncate plans when later waypoints are unreliable.

(a) Success with and without memory, split by whether memory holds a prior success. (b, c) Reliability diagrams before and after calibration.
(a) Success on the same starts with and without memory. (b, c) Predicted waypoint confidence vs. actual reach rate, over 8,458 executed waypoints, before and after calibration.

Downstream policy learning

On six tasks we train Diffusion Policy (fine-tuned from RoboCasa365 and from scratch) and ACT on paired CAPEX and human datasets from identical initial states. At 20k steps, CAPEX-trained policies approach human-trained performance. With longer training the gap largely closes for policies trained from scratch: at 100k steps, ACT reaches 0.50 vs. 0.50 and scratch Diffusion Policy 0.45 vs. 0.47.

Downstream success vs. training steps for DP fine-tuned, DP from scratch, and ACT, comparing CAPEX and human demonstrations.
Downstream success vs. training budget, averaged over five tasks. The fine-tuned Diffusion Policy starts from a checkpoint pretrained on human data, which favors the human condition.

Simulation policy rollouts (Diffusion Policy trained from scratch on CAPEX demonstrations only)

Close Fridge · 0.79 success at 20k steps

Close Toaster Oven Door · 0.51

Turn On Microwave · 0.27

BibTeX

@inproceedings{anonymous2027capex,
  title     = {{CAPEX}: Efficiently Distilling Foundation Model Behavior into Deployable
               Robot Policies through Experience-Adaptive Reasoning},
  author    = {Anonymous},
  booktitle = {Submitted to the International Conference on Learning Representations (ICLR)},
  year      = {2027},
  note      = {Under review}
}