VideoPhysEdit Physical Counterfactual Video Editing
via Rigid-Body Physical Scene Reconstruction

Change not only how a scene looks, but also what happens after a physical edit.

Fudan University emblem Institute of Trustworthy Embodied AI, Fudan University
Choose an edit, click an object, and watch how the scene changes.

Loading the interactive preview…

Abstract

Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods.

A new task

Physical Counterfactual Video Editing

We formulate physical counterfactual video editing (PCVE) as a new, unified task for physical interventions and their downstream consequences. Given a source video, a physical edit, and its execution frame, the goal is to generate a counterfactual video that preserves the factual history before the intervention and depicts the resulting motion and interactions afterward.

Vcf= 𝓕(Vsrc,e,te)
Vsrc
Source video
Vcf
Counterfactual video
e
Physical edit instruction
te
Execution frame
PCVE-RigidBench domino examples: the source sequence and counterfactual sequences after increasing the second domino’s mass, removing the second domino, or removing the first domino at frame 24.
Examples from PCVE-RigidBench show how physical interventions change subsequent motion and interactions.

A new method

Explicit Physical Reasoning Guides Visual Generation

A training-free pipeline for physical counterfactual video editing in rigid-body scenes.

Figure 2. The seven stages of VideoPhysEdit, from object identification and tracking through physical scene reconstruction, physical intervention, and counterfactual video generation.
Overview of VideoPhysEdit. We reconstruct an executable physical scene from the source video, apply the physical intervention, and use the simulated counterfactual trajectories and an edited reference image to guide counterfactual video generation. Open original PDF

A new benchmark

PCVE-RigidBench

20 synthetic rigid-body scenes and 129 editing tasks, with paired source and counterfactual target videos and physical ground truth.

0.376Physical Edit ScoreThe only positive PES among the compared methods.
54.0% Trajectory ErrorRelative to the strongest competing method.
Comparison on PCVE-RigidBench. Bold marks the best result among methods.
MethodPhysical Edit AccuracyVisual Fidelity
PES ↑TE ↓Mask IoU ↑PSNR ↑SSIM ↑LPIPS ↓CLIP ↑FVD ↓
VACEβˆ’0.042146.260.27314.060.7280.4470.7951184.37
Dittoβˆ’0.120149.940.26422.070.8120.2060.850551.14
MiniMax H3βˆ’0.096152.000.25024.870.8700.1270.906246.45
Seedance 2.5βˆ’0.087144.990.23128.850.9280.0800.932246.67
No edit0.000143.130.28931.230.9740.0360.957249.68
VideoPhysEdit0.37666.700.42127.510.9250.1040.929182.46

Comparison with Baselines

Synthetic Videos

Physical edits and their downstream consequences on PCVE-RigidBench.

0.0 s

Real Videos

Physical edits in real-world scenes.

0.0 s

Object Removal

Additional comparisons with VOID on object removal tasks.

0.0 s

BibTeX

Citation details will be added when the paper is available.