Hongyu Liu

PhD Candidate

hongyu/hongyu_profile_new.webp

Video Generation · World Models · Spatial Intelligence

Video Generation World Models Spatial Intelligence Digital Avatars

I am a Ph.D. student in the Department of Computer Science and Engineering at The Hong Kong University of Science and Technology [HKUST], where I am fortunate to be advised by [Prof. Qifeng Chen].

My research centers on video generation and generative world models, human-centered multimodal spatial intelligence, and image generation and editing. I study how generative systems can understand, predict, and synthesize visual worlds across space and time, with a particular interest in realistic, controllable, and interactive digital humans and avatars.

I am particularly interested in real-time and long-form video generation, as well as interactive world models that maintain persistent state, spatial consistency, and controllable scene evolution over extended horizons. My goal is to build efficient generative systems that can perceive actions, model future dynamics, and produce coherent, responsive visual experiences for both open-ended worlds and digital humans.

Beyond research, I am deeply passionate about fishing. I hope to one day travel across China to fish in its rivers, lakes, and coastal waters, using the journey to reconnect with nature, explore local cultures, and find ideas outside the lab.

News

Jul 24, 2026 LiveLight: Real-time Streaming Video Relighting with Interactive Control is accepted to ACM Transactions on Graphics (TOG) 2026.
Jul 09, 2026 We released OPSD-V, an on-policy self-distillation project for post-training few-step autoregressive video generators.
Jun 30, 2026 LACON: Training Text-to-Image Model from Uncurated Data is accepted to ECCV 2026.
Jun 02, 2026 We released DyaPlex, NVIDIA’s full-duplex speech-motion model for dyadic interaction.
May 26, 2026 We released the LongCat-Video-Avatar 1.5 Technical Report, an open-source audio-driven avatar video generation system.

Recent Highlights

Highlighted Research Topics

Three connected directions: image generation and editing, human-centered multimodal spatial intelligence, and video generation with generative world models.

All publications