Previously, I got my MPhil degree from HKUST working with Prof. Ying-Cong Chen and my BEng degree from Zhejiang University.
I spent a wonderful summer at Tencent Hunyuan as a research scientist intern.
Feel free to reach out if you have anything to discuss!
LDR learns "how the future evolves" rather than "what the future is", making it the first video world model that captures the underlying dynamics purely from pixels and extrapolates them beyond the training distribution.
Built on Self Forcing (trained on only 5s clips), Rolling Sink effectively scales the autoregressive video synthesis to ultra-long durations (e.g., 5-30min) at test time, with consistent subjects, stable colors, and smooth motions.
Powered by large-scale curated data and a sphere-aware ViT, DA2 predicts dense distance from a single 360° panorama in an end-to-end manner, with remarkable geometric fidelity and strong zero-shot generalization.
Lotus-2 is an advanced monocular geometric estimator built upon FLUX. By effectively analyzing the DiT-based rectified-flow formulation, Lotus-2 achieves SoTA performance while producing significantly finer details.
Based on Stable Diffusion, Lotus delivers SoTA performance on monocular depth & normal estimation with a simple yet effective fine-tuning protocol that better fits the pre-trained visual prior for dense prediction.
DisEnvisioner effectively identifies and enhances the subject-essential features while filtering out other irrelevant ones, enabling exceptional image customization in a tuning-free manner with only a single image.
DIScene is capable of generating complex, high-fidelity 3D scene with decoupled objects and clear interactions, through a learnable scene graph and hybrid Mesh-Gaussian representation.