CoDrive: Cross-Vehicle Consistent World Video Generation with Precise Trajectory Control for Cooperative Driving

Anonymous supplementary material for double-blind review.

Generated multi-view video (left) and the GT BEV trajectory visualization (right).

Abstract

Real-world driving is inherently multi-agent, yet existing driving world models mainly generate videos from a single vehicle and cannot guarantee consistency across different agents observing the same environment. We present CoDrive, a cross-vehicle, multi-view driving video generation framework that jointly generates observations of vehicles sharing the same dynamic scene with precise camera-trajectory control. CoDrive interleaves local self-attention for intra-agent multi-view modeling with global self-attention for information exchange across vehicles. Camera trajectories are represented in a shared world coordinate system and injected into attention through projective relative positional encoding. We further introduce CoDrive-Bench to evaluate trajectory controllability, scene geometry consistency, and instance-level consistency.

Method

CoDrive extends a pretrained latent video diffusion transformer to jointly model multiple agents and views while preserving explicit camera control.

CoDrive model architecture and progressive training framework
CoDrive jointly denoises multi-agent, multi-view latents with explicit projective camera geometry.
01

Independent encoding

Each view is encoded independently, then concatenated in latent space for joint reasoning without VAE seam artifacts.

02

Local attention

Models spatial-temporal dependencies among the three cameras belonging to each individual vehicle.

03

Global attention

Exchanges information across vehicles so overlapping observations reconcile into one shared scene.

04

Projective geometry

PRoPE injects calibrated relative camera geometry directly into attention for precise trajectory control.

Built on
HunyuanVideo-1.5

Two vehicles × three views · shared reference frame · progressive real + synthetic training

Qualitative Results

Each video contains three forward-facing views from Agent A in the top row and three views from Agent B in the bottom row. The corresponding BEV plot shows the commanded camera trajectories.

Agent A Agent B Video: generated multi-view observations  ·  BEV: commanded trajectories
01

Real-world urban scene

Multi-agent · 6 views · 6.1 sec

Bird's-eye-view trajectories for result case one
Commanded camera paths
02

Turning pair at sunset

CARLA · Town04 · Clear sunset

Bird's-eye-view trajectories for a synthetic turning pair
Commanded camera paths
03

Turning pair on wet roads

CARLA · Town04 · Wet noon

Bird's-eye-view trajectories for a wet-road turning pair
Commanded camera paths
04

Synthetic urban interaction

Multi-agent · 6 views · 6.1 sec

Bird's-eye-view trajectories for result case four
Commanded camera paths
05

Real-world intersection

Multi-agent · 6 views · 6.1 sec

Bird's-eye-view trajectories for result case five
Commanded camera paths
06

Real-world sequence

Multi-agent · 6 views · 6.2 sec

Bird's-eye-view trajectories for result case six
Commanded camera paths

CoDrive-Bench

CoDrive-Bench evaluates multi-agent driving world models along three dimensions: trajectory controllability, scene geometry consistency, and instance-level consistency.

A

Controllability

Does the video follow the commanded path?

ADE ↓DTW ↓
B

Scene consistency

Do both agents agree on static geometry?

Chamfer ↓Reproj ↓
C

Instance consistency

Is the other vehicle where geometry predicts?

Soft-hit ↑M-Err ↓

Quantitative Results

CoDrive achieves the best performance among the compared methods on all trajectory controllability and cross-agent consistency metrics, while maintaining competitive FVD.

ADE ↓3.05
DTW ↓116
Chamfer ↓1.64
Reproj ↓0.52
Soft-hit ↑0.482
M-Err ↓5.81
200evaluation clips
1,200camera views
400trajectory samples
2real + synthetic domains