Egocentric World Models
Learning how dynamic environments evolve from first-person video, and how an agent can predict and act within them.
Computer Vision · Generative AI
Postdoctoral Researcher, Computer Vision Department
Mohamed bin Zayed University of Artificial Intelligence
I am a postdoctoral researcher in the Computer Vision Department at MBZUAI, advised by Prof. Hao Li. My research centers on human-centric generative modeling across 2D images, 3D content, and video, with an emphasis on creating realistic, controllable, and interactive representations of humans and their environments.
Recently, I have focused on world models, particularly egocentric world models and world action models that understand, predict, and interact with dynamic environments from a first-person perspective. These models have broad potential for virtual reality, embodied intelligence, and interactive human-centered simulation.
Previously, I was a visiting Ph.D. researcher at the Robotics Institute, Carnegie Mellon University, advised by Prof. Fernando De la Torre. I completed my doctoral studies at Sun Yat-sen University, supervised by Prof. Xiaodan Liang.
My research connects generative modeling, 3D vision, and embodied perception.
Learning how dynamic environments evolve from first-person video, and how an agent can predict and act within them.
Controllable synthesis of identity, pose, motion, and interaction across images and video.
High-fidelity 3D representations for avatars, virtual try-on, immersive media, and interactive simulation.
Recent research and career updates.
HairOrbit is accepted by ECCV 2026.
GRPO-Guard is accepted by CVPR 2026.
I joined the Computer Vision Department at MBZUAI as a postdoctoral researcher, advised by Prof. Hao Li.
G3D-VTON is accepted by TPAMI 2025.
Robust-MVTON is accepted by CVPR 2025.
I completed my doctoral studies at Sun Yat-sen University and was named an Honors Graduate.
ViTon-GUN is accepted by TVCG 2025.
DreamVTON is accepted by ACM MM 2024.
GUESS is accepted by TVCG 2024.
B2A-HDM is accepted by AAAI 2024.
I began a visiting research period at the Robotics Institute of Carnegie Mellon University.
GP-VTON is accepted by CVPR 2023.
3D-GCL is accepted by NeurIPS 2022.
ARMANI is accepted by ACM MM 2022.
wFlow is accepted by CVPR 2022.
PASTA-GAN is accepted by NeurIPS 2021.
CPF-Net is accepted by TIP 2021.
M3D-VTON is accepted by ICCV 2021.
WAS-VTON is accepted by ACM MM 2021.
Selected peer-reviewed publications. My name is highlighted; * indicates equal contribution.
European Conference on Computer Vision (ECCV), 2026
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
IEEE Transactions on Visualization and Computer Graphics (TVCG), 2025
ACM International Conference on Multimedia (ACM MM), 2024
AAAI Conference on Artificial Intelligence (AAAI), 2024
IEEE Transactions on Visualization and Computer Graphics (TVCG), 2024
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
Advances in Neural Information Processing Systems (NeurIPS), 2022
ACM International Conference on Multimedia (ACM MM), 2022
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
Advances in Neural Information Processing Systems (NeurIPS), 2021
ACM International Conference on Multimedia (ACM MM), 2021
IEEE Transactions on Image Processing (TIP), 2021
IEEE/CVF International Conference on Computer Vision (ICCV), 2021
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
Academic appointments, training, and industry research experience.
Postdoctoral Researcher, Computer Vision Department
Advisor: Prof. Hao Li
Visiting Ph.D., Robotics Institute
Advisor: Prof. Fernando De la Torre
Ph.D. studies, Intelligent Systems Engineering
Advisor: Prof. Xiaodan Liang
M.S., Computer Science and Engineering
Advisor: Prof. Jianhuang Lai
B.S., Computer Science and Engineering
Advisor: Prof. Xiaohua Xie
Research Intern, Intelligence Creation Platform
High-fidelity and 3D virtual try-on
Research Intern, Artificial Intelligence Platform Department
Cross-modal human motion synthesis
Research Intern, Intelligence Creation Platform
High-fidelity 2D garment-to-person virtual try-on
Conferences: CVPR, ICCV, ECCV, NeurIPS
Journal: International Journal of Computer Vision (IJCV)
Organizer, CVPR 2020 Workshop on Human-centric Image/Video Synthesis.
Teaching Assistant, Artificial Intelligence Experiment, Sun Yat-sen University, 2020–2021.
VALSE Webinar on scalable unpaired virtual try-on, Jan. 2022.
Honors Graduate, Sun Yat-sen University, May 2025.