ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

Anonymous Project Page

Rebuttal updates below.

Explore trajectory generation with ODEWorld. Continuous-time dynamics enable denser temporal sampling beyond the training data while maintaining visually plausible rollouts. Long-horizon trajectories can also be generated with replanning. Note that replanning points are realigned for visualization clarity. Drag the timeline to explore any frame.

Results on Agilex

Select a case to explore

Conditioning initial state
initial state (prev. generated)
Goal state
goal image
→ ODE
Rollout
Task Instruction

latent trajectory (PCA)
→ Decode
prediction — overall progress: 0.00%

Results on AgiBot

Select a case to explore

Conditioning initial state
initial state (prev. generated)
Goal state
goal image
→ ODE
Rollout
Task Instruction

latent trajectory (PCA)
→ Decode
prediction — overall progress: 0.00%

Results on Libero

Select a case to explore

Conditioning initial state
initial state (prev. generated)
Goal state
goal image
→ ODE
Rollout
Task Instruction

latent trajectory (PCA)
→ Decode
prediction — overall progress: 0.00%

Additional Comparison vs VPP

Compared with Video Prediction Policy (VPP), a strong pixel-level discrete-time predictor, ODEWorld achieves higher video generation quality and stronger LIBERO-LONG policy performance under the same setting.

Video generation quality

Method Short Horizon @16 frames Long Horizon @64 frames
PSNR ↑ LPIPS ↓ PSNR ↑ LPIPS ↓
VPP 17.01 0.275 13.43 0.530
ODEWorld (Ours) 20.53 0.109 19.46 0.134

LIBERO-LONG policy performance

Method Avg. Suc ↑
VPP 81.0
ODEWorld (Sequential subgoals) 83.6

Real-Robot Control on Agilex

We evaluate ODEWorld on a bi-manual Agilex robot with X-VLA as the policy backbone, augmented by sequential subgoals planned by ODEWorld. Across four tasks, average success rate improves from 55% to 80%.

Method Shrimp Mouse Pen Bi-manual Average
X-VLA 80% 60% 30% 50% 55%
X-VLA + ODEWorld
(Sequential subgoals)
90% 80% 70% 80% 80%

Ablations on Key Design Choices

Additional ablations on the vision backbone, JVP supervision, dynamical decoupling, and latent token count.

Feature extractor

We compare DINOv2 with SigLIP 2 and LIV. ODEWorld works robustly with alternative vision backbones, while DINOv2 remains best on both short- and long-horizon prediction.

Method Short Horizon @16 frames Long Horizon @64 frames
PSNR ↑ LPIPS ↓ PSNR ↑ LPIPS ↓
LIV 19.90 0.158 18.25 0.237
SigLIP 2 20.38 0.133 19.21 0.168
DINOv2 20.53 0.109 19.46 0.134

First-order JVP supervision vs consistency loss

Replacing direct first-order JVP supervision with a traditional consistency loss (indirect zero-order reconstruction of future frames via ODE planning) leads to optimization difficulties, much weaker short-horizon prediction, and visually implausible long-horizon rollouts.

Method PSNR ↑ LPIPS ↓
consistency-loss supervision 12.71 0.751
ours (JVP) 20.53 0.109

Dynamical representation decoupling

Without initial-state-conditioned decoupling, the model uses a high-dimensional non-decoupled latent (\(z_t' \in \mathbb{R}^{256\times 768}\)). Decoupling substantially improves inference speed and long-horizon LPIPS by letting the velocity model focus on temporal evolution rather than static appearance.

Method Short Horizon @16 frames Long Horizon @64 frames
PSNR ↑ LPIPS ↓ FPS ↑ PSNR ↑ LPIPS ↓ FPS ↑
w/o decouple 20.65 0.127 7.301 19.53 0.212 2.29
Ours 20.53 0.109 33.67 19.46 0.134 13.83

Number of latent tokens

Increasing tokens from 1 to 4 yields comparable performance, indicating a single latent token is already sufficient to capture the underlying dynamics with a compact representation.

# Tokens Short Horizon @16 frames Long Horizon @64 frames
PSNR ↑ LPIPS ↓ PSNR ↑ LPIPS ↓
1 20.53 0.109 19.46 0.134
4 20.99 0.114 19.68 0.148

Representation Collapse Analysis

Following RankMe, we measure the effective rank of mean-centered representations. ODEWorld’s 768-d dynamics latent maintains effective rank above 400 across planning horizons, higher than pooled V-JEPA 2 features (~200 / 1024) and DINOv2 CLS (~376 / 768).

Centered RankMe of ODEWorld (effective rank / latent dim)

Horizon 1 4 8 16 32
Centered RankMe 417.2/768 413.9/768 422.1/768 433.1/768 439.5/768

Centered RankMe of V-JEPA 2 (effective rank / latent dim)

Horizon 1 4 8 16 32
Centered RankMe 200.4/1024 201.5/1024 202.9/1024 205.5/1024 208.4/1024

Centered RankMe of DINOv2 (effective rank / latent dim)

DINOv2 feature
Centered RankMe 376.1/768

End of rebuttal updates.

Physical-Time World Modeling

PT-Flow models dynamics as a continuous latent velocity field in physical time, rather than discrete next-step prediction.

  • ODE-based future prediction in compressed latent space.
  • Arbitrary-time queries, and even backward prediction.
  • Compact latents for efficient rollouts.

Bidirectional & Any-Resolution Generation

Forward and backward prediction

Bidirectional generation. Continuous-time formulation enables both forward and backward prediction with physically plausible motion.

Any-resolution generation

Any-resolution temporal generation. ODEWorld supports flexible temporal speed and can recover missing intermediate observations.

Critical Designs

The core idea of PT-Flow is to capture the underlying dynamics as a continuous velocity field within a proper latent space. Let \( s_t \in \mathbb{R}^{d_s} \) denote the state input at physical time \( t \), and \( z_t \in \mathbb{R}^{d_z} \) be the latent representation of \( s_t \). Here, \( t \) represents the physical time elapsed after starting at the initial state \( s_0 \). The dynamics are then modeled by learning a time-dependent velocity field via an ODE formulation.

\[ \frac{dz_t}{dt} = v_{\theta}(z_t, t; z_0, c) \implies z_T = z_0 + \int_0^T v_{\theta}(z_t, t; z_0, c)\,dt \tag{1} \]

To enable such continuous-time modeling, PT-Flow features two key designs:

(1) Dynamical Representation Decoupling. PT-Flow isolates time-varying dynamics from static content by learning initial-state-conditioned dynamics encoder \( f_{\text{dyn}} \) and decoder \( g_{\text{dyn}} \), enabling the latent to focus on dynamics while static context is handled through conditioning.

\[ \mathcal{L}_{\text{dyn-recon}} = \mathbb{E}_{s_0, s_t \sim \mathcal{D}} \left\| g_{\text{dyn}}(f_{\text{dyn}}(s_t; s_0); s_0) - s_t \right\|^2 \tag{2} \]

(2) Direct First-Order Supervision. PT-Flow directly supervises the latent velocity field with the approximated ground truth latent velocities through an elegant Jacobian-vector product (JVP) projection. Note that the time derivative of the latent satisfies:

\[ \dot{z}_t = \frac{d z_t}{d t} = \frac{d f_{\text{dyn}}(s_t; s_0)}{d t} = \frac{\partial f_{\text{dyn}}(s_t; s_0)}{\partial s_t} \cdot \frac{d s_t}{d t} + \bcancel{\frac{\partial f_{\text{dyn}}(s_t; s_0)}{\partial s_0} \cdot \frac{d s_0}{d t}} = \frac{\partial f_{\text{dyn}}(s_t; s_0)}{\partial s_t} \cdot \dot{s}_t \tag{3} \]

As \( s_0 \) is time-invariant, hence \( \frac{d s_0}{d t} = 0 \) and the related term vanishes. This enables us to introduce the following learning objective for the latent velocity field \( v_{\theta} \), which directly supervises it with the approximated ground-truth latent velocities \( \dot{z}_t \).

\[ \mathcal{L}_v = \mathbb{E}_{s_0, s_t \sim \mathcal{D}, t} \left\| v_\theta(z_t, t; z_0, c) - \texttt{sg}\left( \text{JVP}(f_{\text{dyn}}(s_t; s_0), \dot{s}_t) \right) \right\|^2 \tag{4} \]

ODEWorld: Latent World Modeling via PT-Flow

We instantiate the PT-Flow paradigm as ODEWorld, a highly versatile and efficient continuous-time predictive architecture for both high-fidelity video generation and effective robotic policy learning. Apart from building PT-Flow, we first employ a frozen DINOv2 encoder as our vision backbone \( f_{\text{obs}} \) and a dedicated image decoder \( g_{\text{obs}} \) to reconstruct DINO features back to images \( \hat{x}_t \) for video generation. We further embed PT-Flow within the DINO feature space.

ODEWorld framework

In our experiments, we find that a single token is sufficient to model \( z_t \) while achieving strong performance. Specifically, \( z_t \in \mathbb{R}^{1 \times 768} \) is highly compact compared to \( s_t \in \mathbb{R}^{16 \times 16 \times 768} \) in the original DINO space. More importantly, this result further supports the key insight of latent world modeling: raw pixels contain substantial redundancy, while the underlying dynamics can be effectively captured by a much smaller latent representation.

Latent Smoothness Experiments

Latent space analysis.We compare the latent space of ODEWorld against the raw DINO feature space and the latent space of a discrete next-step predictive baseline. PCA visualizations of latent trajectories across multiple LIBERO tasks show that ODEWorld produces the smoothest latent trajectories. Crucially, while improving dynamical smoothness, ODEWorld preserves a manifold topology remarkably similar to that of the original DINO feature space.

Latent smoothness across different latent spaces

Ablation on Dynamical Decoupling. We observe that Decoupling is critical for continuous dynamics modeling. Without separating static context from the state space, the learned latent trajectories become noisy and unstable. Conditioning on \( s_0 \) enables the latent representation to focus on dynamics, leading to smooth and meaningful latent flows.

Ablation on Dynamical Decoupling.

Quantitative Comparisons for Video Generation

Evaluation of Quality and Efficiency. ODEWorld consistently outperforms baselines in both pixel-level and perceptual metrics across short- and long-term horizons, while achieving higher inference speed and a smaller model size.

Method Short Horizon @16 frames Long Horizon @64 frames Params. ↓
PSNR ↑ LPIPS ↓ FPS ↑ PSNR ↑ LPIPS ↓ FPS ↑
LDP 16.33 0.489 0.906 16.27 0.461 0.253 110.9 M
V-JEPA 2 17.60 0.157 8.118 16.47 0.166 1.615 452.34 M
ODEWorld (Ours) 20.53 0.109 33.67 19.46 0.134 13.83 86.08 M

Comparison of Generated Samples. ODEWorld enables long-horizon video generation, producing temporally coherent and physically plausible videos that closely match the ground-truth observations. In contrast, baseline methods often suffer from blurry visual predictions and hallucinations in long-horizon generation.

LIBERO qualitative comparison among ODEWorld, LDP, and V-JEPA 2

Quantitative Comparisons for Policy Learning

ODEWorld for Policy Learning. We assess how ODEWorld can empower downstream policy learning by inferring latent subgoals as guidance. We instantiate three guidance paradigms: (1) velocity-conditioned policy \( \pi(a \mid z, c, v) \), (2) single-subgoal policy \( \pi(a \mid z, c, \hat{z}_{\tau}) \), where \( \tau = 0.25 \), (3) sequential-subgoal policy \( \pi(a \mid z, c, \{\hat{z}_{\tau_i}\}_{i=1}^{n}) \), where \( n = 5 \), \( \tau_i = 0.05 i \), where subgoals are generated via ODE integration. \( c \) denotes image goal and language instruction. Results show that all ODEWorld variants outperform baselines on LIBERO-LONG tasks, with the sequential-subgoal paradigm achieving the highest success rate. This performance gain demonstrates the high-fidelity temporal consistency of ODEWorld rollouts. and such consistent sequential subgoal guidance effectively benefits downstream policy learning.

Method 1 2 3 4 5 6 7 8 9 10 Avg. Suc ↑
GLCBC 83.3 93.3 83.3 93.3 76.6 63.3 73.3 83.3 50.0 80.0 78.0
SuSIE 83.3 63.3 96.6 100.0 83.3 83.3 83.3 39.9 53.3 76.6 76.3
Seer 88.3 90.0 91.6 81.6 85.0 65.0 86.6 80.0 51.6 66.6 78.6
ODEWorld
(Velocity)
80.0 93.3 76.6 100.0 73.3 80.0 83.3 86.6 66.6 83.3 82.3
ODEWorld
(Single subgoal)
86.6 86.6 86.6 96.6 86.6 86.6 80.0 83.3 46.6 86.6 82.6
ODEWorld
(Sequential subgoals)
90.0 96.6 90.0 100.0 73.3 90.0 76.6 90.0 60.0 80.0 83.6