ONE-STEP WORLD MODELING RESEARCH · 2026

Beckmann World Models

Direct Terminal-Map Learning for
One-Step Action-Conditioned Prediction

PushT

Planar manipulation
Ground truthBWM + spatial

Robomimic Can

Robot manipulation
Ground truthBWM + spatial
4 observed + 60 predicted frames · 10 fps playback

Video comparisons

01 / 64Observed

Selected example 02 · PushT

SELECT AN EXAMPLE

Selected final-checkpoint samples · local baselines at 5 epochs

Prediction quality

PushT LPIPS ↓
PushT 64-frame LPIPS by method, with exact values in the table below.
Robomimic Can LPIPS ↓
Robomimic Can 64-frame LPIPS by method, with exact values in the table below.

Local baselines: 5 epochs. BWM: staged training; unequal cumulative compute. Bold = best within each table.

MSE: RGB [−1, 1]. Timing: A800 · batch 1 · FP32 · TF32 off.

Released checkpoints

64 frames · reference only
DatasetModelMSE ↓Moving MSE ↓LPIPS ↓Calls / chunkScope

Policy evaluation & planning

ModelPolicy evaluationGPC-RANK planning
Pearson r ↑95% CIIoU MAE ↓Score ↑Δ vs. DriftWorld [95% CI]

Policy: 7 policies × 300 paired trials. Planning: 50 fixed seeds. CI: evaluation bootstrap. AVDC: unavailable.

Policy correlation Pearson r · 95% CI
Policy correlation with 95% confidence intervals. BWM plus spatial: 0.9462 [0.805, 0.993].
Planning difference Δ vs. DriftWorld · 95% CI
Paired planning differences relative to DriftWorld. BWM plus spatial: −0.001 [−0.055, +0.054]; its interval includes zero.

Action conditioning

True actionsSwapped actions
PushT
Moving-region MSE for true and swapped actions on PushT. BWM plus spatial: 0.0642 versus 0.7089, an 11.0-fold increase.
Robomimic Can
Moving-region MSE for true and swapped actions on Can. BWM plus spatial: 0.0073 versus 0.0451, a 6.2-fold increase.

64 validation windows · one chunk · real history. Action sensitivity, not a counterfactual-accuracy test.