Haodong Li (李浩东)

I'm a PhD student at UC San Diego working with Prof. Manmohan Chandraker. I'm also a research scientist intern at Adobe Research working with Dr. Shaoteng Liu and Dr. Zhe Lin. I'm interested in generative world modeling in both pixels and 3D.

Previously, I got my MPhil degree from HKUST working with Prof. Ying-Cong Chen and my BEng degree from Zhejiang University. I spent a wonderful summer at Tencent Hunyuan as a research scientist intern. Feel free to reach out if you have anything to discuss!

Email: hal211@ucsd.edu, haodongl@adobe.com
Wechat: lihaodong-2000

Google Scholar  /  CV  /  Github  /  Twitter (X)  /  Linkedin  /  Huggingface

profile photo

Selected Publications

Both authors contributed equally.
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Haodong Li, Shaoteng Liu, Tianyu Wang, Chongjian Ge, Sihui Ji, Jiahan Zhang, Xin Lin, Haolin Lu, Zhe Lin, Manmohan Chandraker

arXiv 2026
arXiv / Paper / Project Page / Code GitHub stars / Model / Data

LDR learns "how the future evolves" rather than "what the future is", making it the first video world model that captures the underlying dynamics purely from pixels and extrapolates them beyond the training distribution.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
Haodong Li, Shaoteng Liu, Zhe Lin, Manmohan Chandraker

arXiv 2026
arXiv / Paper / Project Page / Code GitHub stars / Demo / Gallery / Slides

Built on Self Forcing (trained on only 5s clips), Rolling Sink effectively scales the autoregressive video synthesis to ultra-long durations (e.g., 5-30min) at test time, with consistent subjects, stable colors, and smooth motions.

DA2: Depth Anything in Any Direction
Haodong Li, Wangguangdong Zheng, Jing He, Yuhao Liu, Xin Lin, Xin Yang, Ying-Cong Chen, Chunchao Guo

ICLR 2026
arXiv / Paper / Project Page / Code GitHub stars / Demo / Data / Slides

Powered by large-scale curated data and a sphere-aware ViT, DA2 predicts dense distance from a single 360° panorama in an end-to-end manner, with remarkable geometric fidelity and strong zero-shot generalization.

Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
Jing He, Haodong Li, Mingzhi Sheng, Ying-Cong Chen

arXiv 2025
arXiv / Paper / Project Page / Code GitHub stars / Demo (D) / Demo (N)

Lotus-2 is an advanced monocular geometric estimator built upon FLUX. By effectively analyzing the DiT-based rectified-flow formulation, Lotus-2 achieves SoTA performance while producing significantly finer details.

Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
Jing He, Haodong Li, Wei Yin, Yixun Liang, Leheng Li, Kaiqiang Zhou, Hongbo Zhang, Bingbing Liu, Ying-Cong Chen

ICLR 2025
arXiv / Paper / Project Page / Code GitHub stars / Demo (D) / Demo (N) / ComfyUI

Based on Stable Diffusion, Lotus delivers SoTA performance on monocular depth & normal estimation with a simple yet effective fine-tuning protocol that better fits the pre-trained visual prior for dense prediction.

DisEnvisioner: Disentangled and Enriched Visual Prompt for Image Customization
Jing He, Haodong Li, Yongzhe Hu, Guibao Shen, Yingjie Cai, Weichao Qiu, Ying-Cong Chen

ICLR 2025
arXiv / Paper / Project Page / Code GitHub stars / Demo

DisEnvisioner effectively identifies and enhances the subject-essential features while filtering out other irrelevant ones, enabling exceptional image customization in a tuning-free manner with only a single image.

DIScene: Object Decoupling and Interaction Modeling for Complex Scene Generation
Xiao-Lei Li, Haodong Li, Hao-Xiang Chen, Tai-Jiang Mu, Shi-Min Hu

SIGGRAPH Asia 2024
Paper / Video

DIScene is capable of generating complex, high-fidelity 3D scene with decoupled objects and clear interactions, through a learnable scene graph and hybrid Mesh-Gaussian representation.

Academic Service

Reviewer: CVPR 2025 (1), ICLR 2026 (5), CVPR 2026 (2), ICML 2026 (6), SIGGRAPH 2026 (1), ECCV 2026 (2), NeurIPS 2026 (4), WACV 2027 (1), AAAI 2027 (4).