LDR learns "how the future evolves" rather than "what the future is", making it the first video world model that captures the underlying dynamics purely from pixels and extrapolates them beyond the training distribution.
Powered by large-scale curated data and a sphere-aware ViT, DA2 predicts dense distance from a single 360° panorama in an end-to-end manner, with remarkable geometric fidelity and strong zero-shot generalization.
Lotus-2 is an advanced monocular geometric estimator built upon FLUX. By effectively analyzing the DiT-based rectified-flow formulation, Lotus-2 achieves SoTA performance while producing significantly finer details.
Based on Stable Diffusion, Lotus delivers SoTA performance on monocular depth & normal estimation with a simple yet effective fine-tuning protocol that better fits the pre-trained visual prior for dense prediction.
LucidDreamer is a text-to-3D generation framework that distills high-fidelity textures and shapes represented by 3D Gaussians from pre-trained Stable Diffusion with a novel Interval Score Matching objective.