I am a fourth-year PhD student in Artificial Intelligence at HKUST(GZ), advised by Prof. Ying-Cong Chen. I work on controllable and efficient generative models, with a recent focus on long-video generation infrastructure. I am also collaborating closely with Yukang Chen on efficient training and inference systems for video generation.

My research interests center on controllable visual generation, video customization, transparent RGBA video generation, video world models, and practical systems for efficient long-video model training and deployment.

News

  1. We released LongLive-Plug, a once-for-all distillation framework that learns CFG, few-step, and long-context capabilities as reusable LoRAs and plugs them into 54 downstream video models without target training.

  2. We released ReWorld, a real-time interactive world model with long-horizon memory that keeps revisited places consistent over minute-long rollouts.

  3. We released LongLive 2.0, an open-source NVFP4 parallel infrastructure for long-video generation, covering AR training, few-step distillation, multi-shot training and inference, SP acceleration, NVFP4 KV cache, and asynchronous VAE decoding.

  4. We released the survey A Mechanistic View on Video Generation as World Models: State and Dynamics and its companion repository.

Earlier news
  1. Scene Graph Guided Generation, which introduces the Scene Graph Adapter, was accepted to ICCV 2025.

  2. The transparent video generation technology behind TransPixeler has been transferred into Adobe Firefly, enabling transparent-background video generation with foreground and alpha video exports.

  3. TransPixeler was accepted to CVPR 2025.

  4. Motion Inversion was accepted to SIGGRAPH 2025.

  5. Text-Anchored Score Composition was accepted to ECCV 2024.

  6. Selective Diffusion Distillation was accepted to ICCV 2023.

Selected publications

* equal contribution

  1. LongLive-Plug overview: per-task distillation versus once-for-all reusable LoRAs

    LongLive-Plug: Once-for-All Distillation for Video Generation

    Shuai Yang*, Luozhou Wang*, Wei Huang, Zhifei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

    arXiv 2026 New

    Distills CFG, few-step sampling, and long-context error correction once into reusable LoRAs that plug into compatible downstream video models without target training, verified on 54 downstream models.

  2. ReWorld interactive world model rollouts

    ReWorld: An Interactive World Model with Long-Horizon Memory

    Zhifei Chen*, Luozhou Wang*, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu, Yihua Du, Wei Wang, Tianyi Gui, Lianghua Huang, Yingcong Chen

    arXiv 2026 New

    A real-time interactive world model that follows camera actions and remembers revisited places: a fixed-size KV cache with pose-indexed landmark memory keeps minute-long rollouts consistent at 704×1280.

  3. LongLive 2.0 NVFP4 key result overview

    LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

    Yukang Chen*, Luozhou Wang*, Wei Huang*, Shuai Yang*, Bohan Zhang, Yicheng Xiao, Ruihang Chu, Weian Mao, Qixin Hu, Shaoteng Liu, Yuyang Zhao, Huizi Mao, Ying-Cong Chen, Enze Xie, Xiaojuan Qi, Song Han

    arXiv 2026

    A 4-bit long-video generation infrastructure built around NVFP4 and parallelism, supporting the full path from training to deployment and reaching 45.7 FPS on LongLive-2.0-5B.

  4. State Dynamics

    A Mechanistic View on Video Generation as World Models: State and Dynamics

    Luozhou Wang*, Zhifei Chen*, Yihua Du, Dongyu Yan, Wenhang Ge, Guibao Shen, Xinli Xu, Leyi Wu, Man Chen, Tianshuo Xu, Peiran Ren, Xin Tao, Pengfei Wan, Ying-Cong Chen

    arXiv 2026

    A survey that frames video generation as world modeling through state construction and dynamics modeling.

  5. TransPixeler RGBA video preview

    TransPixeler: Advancing Text-to-Video Generation with Transparency

    Luozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang, Zhifei Zhang, He Zhang, Zhe Lin, Yingcong Chen

    CVPR 2025

    The transparent-background video generation technology behind this work has been transferred into Adobe Firefly, where creators can generate foreground videos with alpha-video exports for compositing.

  6. Motion Inversion video customization preview

    Motion Inversion for Video Customization

    Luozhou Wang*, Ziyang Mai*, Guibao Shen, Yixun Liang, Xin Tao, Pengfei Wan, Di Zhang, Yijun Li, Yingcong Chen

    SIGGRAPH 2025

  7. Scene Graph Adapter guidance preview

    Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification

    Guibao Shen*, Luozhou Wang*, Jiantao Lin, Wenhang Ge, Chaozhe Zhang, Xin Tao, Di Zhang, Pengfei Wan, Guangyong Chen, Yijun Li, Ying-cong Chen

    ICCV 2025

    Introduces the Scene Graph Adapter, formerly SG-Adapter, to improve relation control in text-to-image generation.

  8. Text-Anchored Score Composition preview

    Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models

    Luozhou Wang*, Guibao Shen*, Wenhang Ge, Guangyong Chen, Yijun Li, Yingcong Chen

    ECCV 2024

  9. Selective Diffusion Distillation preview

    Not All Steps are Created Equal: Selective Diffusion Distillation for Image Manipulation

    Luozhou Wang*, Shuai Yang*, Shu Liu, Yingcong Chen

    ICCV 2023