Guian Fang

I am a Ph.D. student at Show Lab, National University of Singapore, advised by Prof. Mike Zheng Shou.

Previously, I received my B.Eng. in Intelligent Science and Technology at the School of Intelligent Systems Engineering, Sun Yat-sen University, advised by Xiaodan Liang (梁小丹), co-supervised by Shengcai Liao.

My research centers on generative video models and world models for long-form visual generation, with recent work on video diffusion distillation and embodied intelligence.

Beyond research, I build agentic visual-generation pipelines and runtime infrastructure for coding agents, and I am open to collaboration.

News

  • 2026.06 Two papers — AnyFlow and PAI-Studio — accepted to ECCV 2026.
  • 2026.05 Open-sourced AnyFlow — any-step video diffusion via flow map distillation.
  • 2026.05 Released Claw Orchestrator — a unified runtime for Claude Code, Codex and other coding CLIs.
  • 2026.03 Launched PAI at Utopai Studios — long-form video generation for cinematic storytelling.
  • 2025.07 RealignDiff accepted to IEEE TNNLS — coarse-to-fine semantic re-alignment for diffusion.
  • 2025.06 Launched MikoAI — an ACG (anime / comic / manga) creation tool, powered by FramePrompt.

Publications

Representative works are highlighted. * denotes equal contribution. Full list on Google Scholar · CV.

AnyFlow:
Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, Song Han, Han Cai, Mike Zheng Shou

ECCV, 2026Oral

PAI-Studio:
Cinematic Video Background Replacement with Camera-Aware Motion

Heyuan Gao*, Bangxun Tang*, Yiren Song*, Guian Fang, Zijian He, Jie Yang, Mike Zheng Shou

ECCV, 2026

VEditBench overview — 420 real-world videos, 6 editing tasks, 9 evaluation dimensions

VEditBench:
Holistic Benchmark for Text-Guided Video Editing

Jay Zhangjie Wu, Guian Fang, Dongrong Joe Fu, Vijay Anand Raghava Kanakagiri, Forrest Iandola, Kurt Keutzer, Wynne Hsu, Zhen Dong, Mike Zheng Shou

Preprint, 2025

HumanRefiner research visualization showing human pose refinement

HumanRefiner:
Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance

Guian Fang*, Wenbiao Yan*, Yuanfan Guo*, Jianhua Han, Zutao Jiang, Hang Xu, Shengcai Liao, Xiaodan Liang

ECCV, 2024

T2VScore framework — text alignment and video quality evaluation pipelines

T2VScore:
Towards A Better Metric for Text-to-Video Generation

Jay Zhangjie Wu*, Guian Fang*, Haoning Wu*, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, Yuchao Gu, Rui Zhao, Weisi Lin, Wynne Hsu, Ying Shan, Mike Zheng Shou

arXiv preprint, 2024

ChartThinker framework diagram showing contextual chain-of-thought approach

ChartThinker:
A Contextual Chain-of-Thought Approach to Optimized Chart Summarization

Mengsha Liu, Daoyuan Chen, Yaliang Li, Guian Fang, Ying Shen

LREC-COLING, 2024

RealignDiff framework showing coarse-to-fine semantic re-alignment process

RealignDiff:
Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment

Guian Fang*, Zutao Jiang*, Jianhua Han, Guansong Lu, Hang Xu, Shengcai Liao, Xiaojun Chang, Xiaodan Liang

IEEE TNNLS, 2023

Products & Open Source

Deployed products and open-source systems I've led or co-built — productized research rather than peer-reviewed papers.

Honors & Awards

Scholarships

Competitions

Activities & Services

Conference Reviewer

  • CV: ECCV, CVPR, ICCV
  • ML: NeurIPS, ICLR, ICML
  • AI: AAAI, AISTATS
  • NLP: ACL Rolling Review (ACL, EMNLP, NAACL, EACL)

Workshop Organizer

  • LOVEU Workshop @ CVPR 2024: Long-form Video Understanding Towards Multimodal AI Assistant and Copilot

Teaching Assistant

  • EE3703: Machine Learning with Applications
  • EE4309: Robot Perception
  • EE5106: Advanced Robotics

Acknowledgements

Grateful to the mentors and teams I've worked with along the way.