Independent encoding
Each view is encoded independently, then concatenated in latent space for joint reasoning without VAE seam artifacts.
Real-world driving is inherently multi-agent, yet existing driving world models mainly generate videos from a single vehicle and cannot guarantee consistency across different agents observing the same environment. We present CoDrive, a cross-vehicle, multi-view driving video generation framework that jointly generates observations of vehicles sharing the same dynamic scene with precise camera-trajectory control. CoDrive interleaves local self-attention for intra-agent multi-view modeling with global self-attention for information exchange across vehicles. Camera trajectories are represented in a shared world coordinate system and injected into attention through projective relative positional encoding. We further introduce CoDrive-Bench to evaluate trajectory controllability, scene geometry consistency, and instance-level consistency.
CoDrive extends a pretrained latent video diffusion transformer to jointly model multiple agents and views while preserving explicit camera control.
Each view is encoded independently, then concatenated in latent space for joint reasoning without VAE seam artifacts.
Models spatial-temporal dependencies among the three cameras belonging to each individual vehicle.
Exchanges information across vehicles so overlapping observations reconcile into one shared scene.
PRoPE injects calibrated relative camera geometry directly into attention for precise trajectory control.
Two vehicles × three views · shared reference frame · progressive real + synthetic training
Each video contains three forward-facing views from Agent A in the top row and three views from Agent B in the bottom row. The corresponding BEV plot shows the commanded camera trajectories.
Multi-agent · 6 views · 6.1 sec
CARLA · Town04 · Clear sunset
CARLA · Town04 · Wet noon
Multi-agent · 6 views · 6.1 sec
Multi-agent · 6 views · 6.1 sec
Multi-agent · 6 views · 6.2 sec
CoDrive-Bench evaluates multi-agent driving world models along three dimensions: trajectory controllability, scene geometry consistency, and instance-level consistency.
Controllability
Scene consistency
Instance consistency
CoDrive achieves the best performance among the compared methods on all trajectory controllability and cross-agent consistency metrics, while maintaining competitive FVD.