Video Generation · World Models · Spatial Intelligence
I am a Ph.D. student in the Department of Computer Science and Engineering at The Hong Kong University of Science and Technology [HKUST], where I am fortunate to be advised by [Prof. Qifeng Chen].
My research centers on video generation and generative world models, human-centered multimodal spatial intelligence, and image generation and editing. I study how generative systems can understand, predict, and synthesize visual worlds across space and time, with a particular interest in realistic, controllable, and interactive digital humans and avatars.
I am particularly interested in real-time and long-form video generation, as well as interactive world models that maintain persistent state, spatial consistency, and controllable scene evolution over extended horizons. My goal is to build efficient generative systems that can perceive actions, model future dynamics, and produce coherent, responsive visual experiences for both open-ended worlds and digital humans.
Beyond research, I am deeply passionate about fishing. I hope to one day travel across China to fish in its rivers, lakes, and coastal waters, using the journey to reconnect with nature, explore local cultures, and find ideas outside the lab.
Experience & Affiliations
News
| Jul 24, 2026 | LiveLight: Real-time Streaming Video Relighting with Interactive Control is accepted to ACM Transactions on Graphics (TOG) 2026. |
|---|---|
| Jul 09, 2026 | We released OPSD-V, an on-policy self-distillation project for post-training few-step autoregressive video generators. |
| Jun 30, 2026 | LACON: Training Text-to-Image Model from Uncurated Data is accepted to ECCV 2026. |
| Jun 02, 2026 | We released DyaPlex, NVIDIA’s full-duplex speech-motion model for dyadic interaction. |
| May 26, 2026 | We released the LongCat-Video-Avatar 1.5 Technical Report, an open-source audio-driven avatar video generation system. |
Recent Highlights
OPSD-V
On-policy self-distillation for post-training few-step autoregressive video generators, reducing long-horizon degradation.
LongCat-Video-Avatar 1.5
Open-source audio-driven avatar video generation with stronger lip sync, long-video stability, and practical 8-step inference.
DyaPlex
Full-duplex speech-motion model for dyadic interaction, coupling streaming speech and body motion for responsive digital agents.
Research Map
Highlighted Research Topics
Three connected directions: image generation and editing, human-centered multimodal spatial intelligence, and video generation with generative world models.
Image Generation & Editing
Controllable image synthesis and editing, spatial reconstruction, motion transfer, and content-aware generation.
Human-Centered Spatial Intelligence
Multimodal understanding and generation for digital humans and avatars, with controllable appearance, geometry, motion, and interaction.
Video Generation & World Models
Video generation, interactive world models, temporally coherent scene evolution, and controllable visual simulation.