Efficient world action models

Latent Action as Intention Enables EfficientFuture Imagination for World Action Models

LAWA keeps the benefits of test-time future imagination while replacing expensivefuture-observation generation with compact, executable latent intentions.

Xiang Li1,2,4*, Yupeng Zheng3*†‡, Songen Gu5*, Huailiang Ma6*, Feng Yu7, Xian Nie7, Shanshuai Yuan4,5,
Yujie Zang8, Weize Li8, Shuai Tian3, Moyang Liu9, Ya-Qin Zhang2, Wenchao Ding4†

1CollegeAI, THU2AIR, THU3CASIA4TARS Robotics5FDU6Southeast University7SJTU8NUS9BUAA

* Equal contribution · Corresponding author · Project leader

Conceptual and quantitative comparison of Fast-WAM, Joint-WAM, and LAWA across latency, RoboCasa performance, and scalability.
One compact future interface.Performance, generalization, and latency in a single comparison.
65.6%RoboCasa few-shot success
80.8%RoboCasa full-data success
74.4%LIBERO-Plus zero-shot success
42.9%Lower latency than Joint-WAM

Abstract

Imagine the future without rendering it.

Future imagination need not be discarded for efficiency: LAWA moves it from observation space into a compact latent-action space.

World action models improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Matched implementations show lower generalization for Fast-WAM than future-aware alternatives, especially with scarce demonstrations and out-of-distribution scenarios.

LAWA uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. A discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets; LAWA jointly denoises a continuous latent state anchored to those targets with executable action chunks.

Across RoboCasa, LIBERO-Plus, and real-world tasks, LAWA improves performance over Fast-WAM while preserving the performance level of Joint-WAM at substantially lower inference latency.

The design space

Three ways to connect perception and action.

01 / FAST

Fast-WAM

Removes future prediction at inference, reducing latency but weakening generalization in the matched comparisons.

current viewaction
02 / JOINT

Joint-WAM

Explicitly imagines future observations and feeds them to the action expert, preserving a direct, observation-space representation of task progress. This future awareness comes at the cost of iterative visual generation during inference.

current viewfuture pixelsaction

Method

Joint attention turns latent transitionsinto executable control.

Training connects video, latent-action, and action experts. Inference removes the future-video branch and keeps only the compact intention pathway plus action prediction.

LAWA architecture with video denoising, latent action prediction, action prediction, multi-model joint attention, and structured attention masks for training and inference.
LAWA overview. A structured attention mask blocks future observations while allowing the action expert to use the evolving latent intention.
1

Tokenize transitions

A discrete tokenizer compresses current-to-future visual changes into a sequence of latent actions.

2

Predict intentions

LAWA denoises a continuous latent state anchored to manipulation-centric codebook targets.

3

Execute actions

The action expert attends to the latent intention and jointly denoises executable action chunks.

Latent action tokenizer with DINO encoder, attention and quantization, forward decoder, observation prediction, and mask prediction.

Action-free pre-training

A manipulation-centric tokenizer.

Video reconstruction can favor static appearance over small interaction dynamics. LAWA adds automatically generated mask targets to bias the tokenizer toward hands, manipulators, and interaction regions.

  • Pre-trains on action-free robot and egocentric videos
  • Quantizes temporally structured visual transitions
  • Uses observation and auxiliary mask prediction

Results

A favorable performance–latency trade-off.

The comparisons answer complementary questions: Fast-WAM is the primary performance baseline,while Joint-WAM is the efficiency reference for explicit future generation.

Overall comparison on RoboCasa

Average success rate across 24 tabletop tasks.

MethodParadigmFew-shotFull data
GR00T N1.6VLA47.6
StarVLAVLA48.8
StarVLA-αVLA53.8
TwinBrainVLAVLA54.6
ABot-M0VLA58.3
RLDX-1VLA58.7
JoyAI-RAVLA63.2
DIALVLA58.370.2
Being-H0.7WAM49.2
DiT4DiTWAM50.8
LDA-1BWAM55.4
Fast-WAM†WAM56.076.3
Joint-WAM†WAM64.178.8
LAWA (Ours)WAM65.680.8

LAWA exceeds the matched Fast-WAM by 9.6 points in few-shot and 4.5 points with full data.

LIBERO-Plus zero-shot robustness

Micro-average success across seven perturbation types.

MethodLanguageCameraLightingMotionNoiseTextureLayoutTotal
OpenVLA0.83.523.08.134.815.228.515.6
OpenVLA-OFT56.431.979.588.793.375.874.269.6
UniVLA1.846.269.669.081.021.231.942.9
WorldVLA0.127.941.643.717.110.938.025.0
π013.86.058.885.081.479.068.953.6
π0-FAST65.121.661.073.273.274.468.861.6
Fast-WAM†24.950.776.989.262.358.067.760.0
Joint-WAM†47.265.391.894.857.661.078.970.4
LAWA (Ours)69.266.762.896.264.585.578.674.4

End-to-end latency

Milliseconds per action chunk on one NVIDIA A800 GPU.

Fast-WAM
196.5
Joint-WAM
593.1
LAWA
338.5

LAWA is 42.9% faster than Joint-WAM at comparable success, while substantially outperforming the faster Fast-WAM.

Why the latent sequence matters

RoboCasa full-data success under inference-time perturbations.

Latent inputSuccess rate
Gaussian noise52.2
Temporal shuffle56.4
Unperturbed80.8
Success rates for LAWA and Fast-WAM as action-free egocentric pre-training data increases from 10 to 100 percent.

Scales with action-free video

From 10% to 100% of the video corpus, LAWA gains 3.6 points in full-data and 4.0 points in few-shot training.

Attention heatmaps comparing Fast-WAM and LAWA across five stages of a manipulation rollout.

Focuses on interaction regions

In the illustrative rollout, LAWA follows the manipulated object and task-relevant region rather than diffuse background appearance.

Real-world tasks

From fine-grained assembly to long-horizon manipulation.

Four tasks test complementary requirements on a UFACTORY xArm7: Gear and Battery emphasize assembly precision; Block and Laboratory require multi-stage execution.

Four real-world robot tasks labeled Gear, Battery, Block, and Laboratory with arrows indicating object transfers.
40.0%LAWA average with 25% of demonstrations
67.5%LAWA average with full demonstrations
31.2point gain over Fast-WAM at 25% data
33.8point gain over Fast-WAM at full data

With only 25% of demonstrations—50 trajectories per task—LAWA reaches 40.0% average success, surpassing the 33.8% achieved by Fast-WAM with the full training set. In the full-data setting, LAWA improves by 45 points on both long-horizon tasks.

Robot rollouts

LAWA rollouts in the real world.

Each task includes two representative LAWA rollout recordings captured from different angles. Every video is encoded at 4× its original speed.

Gear assembly

A fine-grained task requiring the robot to position the gear so its teeth mesh with the adjacent gear.

Rollout 1LAWA · 4×
Rollout 2LAWA · 4×

Battery insertion

A fine-grained assembly task in which the battery must be aligned and pressed until fully seated in the slot.

Rollout 1LAWA · 4×
Rollout 2LAWA · 4×

Block manipulation

A long-horizon task that requires multiple object transfers and sustained execution across stages.

Rollout 1LAWA · 4×
Rollout 2LAWA · 4×

Laboratory manipulation

A long-horizon sequence involving multiple containers and transfers across the workspace.

Rollout 1LAWA · 4×
Rollout 2LAWA · 4×

Citation

BibTeX

Citation metadata will be added after the anonymous submission stage.

BibTeX forthcoming.