Computer Vision · Computer Graphics · Generative AI
I am a Ph.D. student in Computer Science and Engineering at HKUST, advised by Prof. Qifeng Chen.
My research spans video generation and generative world models, human-centered multimodal spatial intelligence, and image generation and editing. I currently focus on post-training world models and embodied intelligence.
Outside the lab, I love fishing and hope to explore China’s rivers, lakes, and coastlines, one fishing trip at a time.
Experience & Affiliations
News
| Oct 09, 2026 | Two papers have been accepted to NeurIPS 2026! |
|---|---|
| Oct 06, 2026 | We released WorldSonus, a streaming video-to-audio model for synchronized, controllable spatial sound. |
| Sep 03, 2026 | We released HelixWorld, a real-time interactive audio-visual world model with synchronized video and spatial sound. |
| Jul 24, 2026 | LiveLight: Real-time Streaming Video Relighting with Interactive Control is accepted to ACM Transactions on Graphics (TOG) 2026. |
| Jul 09, 2026 | We released OPSD-V, an on-policy self-distillation project for post-training few-step autoregressive video generators. |
Recent Highlights
OPSD-V
On-policy self-distillation for post-training few-step autoregressive video generators, reducing long-horizon degradation.
HelixWorld
Real-time interactive world generation with synchronized video and spatial sound that follows your viewpoint.
DyaPlex
Full-duplex speech-motion model for dyadic interaction, coupling streaming speech and body motion for responsive digital agents.
Research Map
Highlighted Research Topics
Three connected directions: video generation and generative world models, human-centered multimodal spatial intelligence, and image generation and editing.
Video Generation & World Models
Real-time and long-form video generation, with persistent world state, coherent dynamics, and interactive control.
Human-Centered Spatial Intelligence
Multimodal intelligence for digital humans and avatars, spanning appearance, geometry, motion, and interaction.
Image Generation & Editing
Controllable image generation and editing, with spatial reasoning, motion transfer, and content-aware synthesis.