unlike most work on image/video editing that is not guaranteed to "follow rules" of 3D consistency, our work builds an explicit virtual world that you can move a virtual camera through. we use a statistical learning-based system, but we distill it into a 3D model that follows at least some 3D rules