Flow4R: Unifying 4D Reconstruction and Tracking With Scene Flow
MCML Authors
Abstract
Abstract
Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruction systems assume a rigid world, whereas dynamic tracking frameworks rely on explicit ego-motion estimation or separate object motion models. In this work, we propose Flow4R, a unified framework that treats relative scene flow as the central representation linking 3D structure, camera ego-motion, and dynamic object motion. Given a two-view input, Flow4R employs a shared Vision Transformer to predict a compact, pixel-aligned property set comprising 3D point positions, scene flow, pose weights, and confidence maps. This flow-centric formulation allows local geometry and bidirectional motion to be jointly inferred in a single feedforward pass, eliminating the need for explicit pose regression heads or complex bundle adjustment. By training jointly on static and dynamic datasets, Flow4R achieves state-of-the-art performance on 4D reconstruction and tracking benchmarks, demonstrating the power of the flow-centric formulation for spatiotemporal scene understanding.
inproceedings QZW+26
ECCV 2026
19th European Conference on Computer Vision. Malmö, Sweden, Sep 08-12, 2026. To be published. Preprint available.Authors
S. Qian • G. Zhang • S. Wu • D. CremersLinks
arXivResearch Area
BibTeXKey: QZW+26