Guian Fang

I am a Ph.D. student at Show Lab, National University of Singapore, advised by Prof. Mike Zheng Shou.

Previously, I received my B.Eng. in Intelligent Science and Technology at the School of Intelligent Systems Engineering, Sun Yat-sen University, advised by Xiaodan Liang (梁小丹), co-supervised by Shengcai Liao.

My research centers on generative video models and world models for long-form visual generation, currently on few-step distillation, long-horizon memory, and embodied intelligence.

I also work on agents that produce visual content rather than only consume it, and on runtime infrastructure for coding agents. I am open to collaboration.

News

  • 2026.06 Two papers — AnyFlow and PAI-Studio — accepted to ECCV 2026.
  • 2026.05 Open-sourced AnyFlow with NVIDIA — any-step video diffusion via on-policy distillation.
  • 2026.05 Released Claw Orchestrator — a unified runtime for Claude Code, Codex and other coding CLIs.
  • 2026.03 Launched PAI at Utopai Studios — long-form video generation for cinematic storytelling.
  • 2025.07 RealignDiff accepted to IEEE TNNLS — coarse-to-fine semantic re-alignment for diffusion.

Publications

Representative works are highlighted. * denotes equal contribution. Full list on Google Scholar · CV.

AnyFlow:
Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, Song Han, Han Cai, Mike Zheng Shou

ECCV, 2026Oral

Upstreamed into diffusers · FastVideo · FastGen

PAI-Studio:
Cinematic Video Background Replacement with Camera-Aware Motion

Heyuan Gao*, Bangxun Tang*, Yiren Song*, Guian Fang, Zijian He, Jie Yang, Mike Zheng Shou

ECCV, 2026

Declare, Compile, Look pipeline — plan a typed FigureSpec, compile the geometry deterministically, emit figures, and repair by patching the spec

Declare, Compile, Look:
Coordinate-Free Layout Generation with Vision-in-the-Loop Repair

Guian Fang, Mengsha Liu, Mike Zheng Shou

ECCV Workshop (MDA), 2026Talk

VEditBench overview — 420 real-world videos, 6 editing tasks, 9 evaluation dimensions

VEditBench:
Holistic Benchmark for Text-Guided Video Editing

Jay Zhangjie Wu*, Guian Fang*, Dongrong Joe Fu, Vijay Anand Raghava Kanakagiri, Forrest Iandola, Kurt Keutzer, Wynne Hsu, Zhen Dong, Mike Zheng Shou

Preprint, 2025

HumanRefiner research visualization showing human pose refinement

HumanRefiner:
Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance

Guian Fang*, Wenbiao Yan*, Yuanfan Guo*, Jianhua Han, Zutao Jiang, Hang Xu, Shengcai Liao, Xiaodan Liang

ECCV, 2024

T2VScore framework — text alignment and video quality evaluation pipelines

T2VScore:
Towards A Better Metric for Text-to-Video Generation

Jay Zhangjie Wu*, Guian Fang*, Haoning Wu*, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, Yuchao Gu, Rui Zhao, Weisi Lin, Wynne Hsu, Ying Shan, Mike Zheng Shou

Preprint, 2024

ChartThinker framework diagram showing contextual chain-of-thought approach

ChartThinker:
A Contextual Chain-of-Thought Approach to Optimized Chart Summarization

Mengsha Liu, Daoyuan Chen, Yaliang Li, Guian Fang, Ying Shen

LREC-COLING, 2024

RealignDiff framework showing coarse-to-fine semantic re-alignment process

RealignDiff:
Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment

Guian Fang*, Zutao Jiang*, Jianhua Han, Guansong Lu, Hang Xu, Shengcai Liao, Xiaojun Chang, Xiaodan Liang

IEEE TNNLS, 2023

Products & Open Source

Deployed products and open-source systems I've led or co-built — productized research rather than peer-reviewed papers.

Honors & Awards

Scholarships

Competitions

Activities & Services

Conference Reviewer

  • CV: ECCV, CVPR, ICCV
  • ML: NeurIPS, ICLR, ICML
  • AI: AAAI, AISTATS
  • NLP: ACL Rolling Review (ACL, EMNLP, NAACL, EACL)

Workshop Organizer

  • LOVEU Workshop @ CVPR 2024: Long-form Video Understanding Towards Multimodal AI Assistant and Copilot

Teaching Assistant

  • EE3703: Machine Learning with Applications
  • EE4309: Robot Perception
  • EE5106: Advanced Robotics
  • EE6733: Advanced Topics on Vision and Machine Learning

Acknowledgements

Grateful to the mentors and teams I've worked with along the way.