Video-Point Model (VPM): Learning the 4D World through Point Trajectory Forecasting

Byungwoo Jeon1*, Sunghwan Hong2*, Dohyeon Kim3, Jonghoon Lee1, Jungwoo Park1, Inhee Lee3, Doosan Baek4, Jinil Kim4, Taehyun Lim4, Hanbyul Joo3, Marc Pollefeys2, Saining Xie5†, Jinwoo Shin1†

1KAIST2ETH Zürich3Seoul National University4Physics Sim Lab5New York University

*Equal contribution†Equal advising

TL;DR

Multi-view world models can render every camera convincingly while the cameras disagree about how the scene moves in 3D. VPM (Video-Point Model) jointly forecasts future videos from arbitrary viewpoints and the 3D trajectories of persistent scene points, given posed multi-view histories and a language instruction. Learning where points will move improves both video fidelity and 3D consistency, and the same point supervision also improves robot action learning.

Overview video. Motivation, model, point-guided generation, zero-shot control from predicted tracks, and real-robot rollouts.

Plausible videos are not coherent 3D motion

Reliable multi-view forecasting requires videos from different cameras to describe the same evolving 3D scene. Visual plausibility alone does not establish whether a forecast does.

Cosmos3-Nano / three generated views

Prompt: “The van slows down and makes a right turn at the intersection into the street on its right.”

Each view can look right on its own

Forecasts from arbitrary viewpoints should depict the same future event despite differences in perspective and visibility. Per-view realism does not check that.

Idea: forecast persistent 3D point trajectories together with the videos. Each trajectory follows one physical point across every camera and over time, so supervising it ties appearance changes in all views to a shared description of 3D motion.

Overview figure with three panels: (a) VP-WorldBench and VPM-Data with robot table manipulation and urban dynamics clips shown as 3D point trajectories; (b) the Video-Point Model taking multi-view video histories, camera parameters and a language instruction and producing future multi-view videos, robot actions and persistent 3D point trajectories; (c) benchmark results where VPM (V+T) has the lowest FVD and highest 3D accuracy, and bar charts showing track supervision lowers FVD on Robot and General splits.

Method

VPM builds on the Mixture-of-Transformers backbone of Cosmos 3. It adds camera-conditioned multi-view processing and a trajectory head that is trained jointly with video generation, so point-motion supervision shapes the representation used for visual prediction.

Architecture diagram: a reasoner reads the task instruction; a generator with shared layers processes action, history and noisy-future tokens and outputs action chunks and future videos; features from an intermediate layer feed a point tracker that uses 4D local correlation and Plücker embeddings, refines tracks with spatial, cross-view and temporal refinement for M iterations, and lifts them to a 3D point track.
From posed multi-view histories and a task instruction, VPM forecasts future videos and the trajectories of points queried from observed frames. The point tracker iteratively refines positions and depths before lifting them to 3D. The action-enabled extension also outputs robot action chunks.

VPM-Data & VP-WorldBench

VPM-Data provides 10K simulated training clips that pair synchronized multi-view videos with persistent 3D point trajectories, camera geometry, depth, visibility masks and captions. VP-WorldBench evaluates forecasting on 1.2K held-out clips across simulated tabletop and urban scenes and real-world recordings.

VPM-Data (training): 10K simulation clips, 8,000 tabletop manipulation with robot, objects and clutter, and 2,000 urban clips with humans and vehicles.
VP-WorldBench (evaluation): 1.2K clips. Simulation: 500 tabletop (436 city, plus unseen kitchens) and 500 urban (332 humans, 168 vehicles). Real: 100 table clips with 2 views and 100 indoor clips with 48 views. Hard subsets of 27, 78 and 17 clips.
Red arcs mark hard clips, in which individual cameras lose sight of the target for at least 30 consecutive frames while another camera keeps it in view.

Benchmark clips with ground-truth 3D point tracks

Simulated tabletop manipulation3 views / 20 fps

A Fourier GR-1 T2 pick-and-place episode rendered in Unreal Engine 5 with clutter and distractors. Ground-truth 3D points on the robot's arms and hands and on the manipulated object are projected into each view.

Kinematics-grounded 3D points for real robot videos

For DROID and HRDexDB episodes, points are sampled once on the robot's CAD meshes and moved with the recorded joint states through forward kinematics, so every point keeps its physical identity across time and views. An episode is kept only if its projected CAD mask agrees with a SAM2 robot mask (recall ≥ 0.9).

Columns: raw frame / projected CAD mask / gripper points with trajectories. Rows: two exterior cameras.

Results

VPM (V) is trained on multi-view videos only; VPM (V+T) adds 3D point-trajectory supervision under the same training budget, so the comparison isolates the effect of point supervision. For a fair comparison, no method receives camera parameters for future frames.

VP-WorldBench

Baselines are fine-tuned on the same data; single-view models generate each requested view independently. FVD and PSNR measure video fidelity; MEt3R and δ3Davg measure geometric consistency, where δ3Davg is the average fraction of 3D tracks, extracted from generated videos with MVTracker, that land within depth-adaptive thresholds of the ground truth.

The radar charts require JavaScript.

Point supervision improves robot action learning

The action-enabled VPM (V+T+A) is evaluated on LIBERO, LIBERO-Plus and RoboCasa in simulation, and on three real pick-and-place tasks with a Franka Panda arm in the DROID setup (two exterior cameras and one wrist camera, 24 trials per task).

Success rate (%)

The bar charts require JavaScript.

Methods without a reported number for a benchmark are left out of that panel. Real-world scores give 0.5 for a successful pick-and-place that touches the pile of boxes.

Real-robot rollouts

Each task places the object and its goal on opposite sides of a pile of boxes, so no single camera shows the whole task and the arm must cross the pile without touching it. Clips play at 2× speed.

PnP BananaPut the banana in the basket without touching the boxes
PnP CupPut the cup on the plate without touching the boxes
PnP CubeStack the cube on the red cube without touching the boxes

Zero-shot manipulation from predicted 3D point tracks

Rows from top to bottom: generated video, predicted point trajectory, robot rollout. Columns from left to right: cam0 and cam1 (multi-view), then a 3D visualization.