About Me

I am a ZJU100 Young Professor at Zhejiang University. Previously, I interned at Alibaba DAMO Academy and worked at the Advanced Institute of Information Technology, Peking University.

Research Interests

My research focuses on Large Language Models, Multi-modal Models, and their applications in Embodied Intelligence. I have published over 40 papers at top international AI conferences.

  • LLM Agent: LLM-powered autonomous agents, particularly in autonomous task execution, reasoning, and self-evolution through environmental interaction.

  • Spatial Intelligence: Enhancing models’ spatial cognition capabilities to better perceive, understand, reason, and interact within 3D environments, serving as the embodied brain.

  • Behavioral Intelligence: Developing action generation models, such as Vision-Language-Action models (VLAs), for action generation and imitation, enabling flexible and autonomous navigation and manipulation—serving as the embodied cerebellum.

  • Social Intelligence: Developing emotionally intelligent LLMs that not only excel in reasoning but also understand human intentions, emotions, and goals—enhancing their social capabilities for more empathetic and human-centered interactions.

Lab Goal

🔥 News

📝 Selected Publications

(# indicates corresponding author)

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
Kaixiang Yao, Xu Wang, Miao Pan, Hu Xiyue, Weishi Wang, Daniel Dahlmeier, Jintao Chen, Yongliang Shen, Xuhong Zhang, Wenqi Zhang

arXiv Stars83 Pages 🤗 Model downloads644 Total 🤗 Dataset downloads1,068 Total X 小红书

  • Learn local state transitions and long-horizon spatial reasoning from simulated and real interaction trajectories.
  • Introduce the LSI-108K curriculum and combine supervised fine-tuning with on-policy distillation.

EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
Lizhou Liang, Xinyu Zhong, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Qinfeng Li, Peng Li, Jintao Chen, Xuhong Zhang, Wenqi Zhang

arXiv Embodied-Omni Stars249 Pages 🤗 Dataset downloads1,760 Total X 小红书

  • Benchmark embodied memory through 2,554 interactive episodes across four task families.
  • Introduce Embodied-Memorizer with spatial, event, and scene memories, and train the EMem-8B policy.

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Hongyan Feng, Sunlai Chen, Xuanyu Liu, Miao Pan, Yangfan Xie, Yuxiang Cui, Zhongxiang Zhou, Rong Xiong, Wenqi Zhang, Jianwei Yin, Yueting Zhuang, Xuhong Zhang

arXiv Embodied-Omni Stars249 Pages 🤗 Model downloads79 Total 小红书

  • Bridge visual grounding and 3D navigation through pixel pointing, selective reasoning, and Anchor-Trajectory Memory.
  • Align navigation decisions with Two-Level GRPO and demonstrate zero-shot deployment on a Unitree Go2 quadruped.

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
Yi Pan, Miao Pan, Qi Lu, Jiaming Huang, Man Zhang, Siteng Huang, Xin Li, Jie Zhang, Yongliang Shen, Xuhong Zhang, Wenqi Zhang

arXiv Stars89 Pages 机器之心 具身智能之心TechDaily 小红书

  • Detect execution drift with a lightweight latent-space visual monitor while keeping the VLA backbone frozen.
  • Truncate stale action chunks and trigger corrective replanning for adaptive action horizons.

Show, Don’t Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang

arXiv Stars25 Pages 🤗 Dataset downloads1,643 Total 机器之心 X 小红书

  • Evaluate spatial cognition through protocol-constrained visual answers parsed into comparable metrics.
  • Introduce SpatialGen-Bench with 470 samples across 14 spatial subtasks and four capability levels.
ACL 2026 Findings
GFT method overview

GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
Wangjie Gan, Miao Pan, Linbo Xi, Wenqi Zhang, Jintao Chen, Jianwei Yin, Xuhong Zhang

ACL 2026 arXiv Stars36 青稞AI 小红书

  • Use Group Advantage Learning to derive reward-based supervision from diverse response groups.
  • Stabilize optimization with Dynamic Coefficient Rectification and improve the transition to subsequent reinforcement learning.
ICCV 2025

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
Wenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun, Yongliang Shen, Weiming Lu, Deli Zhao, Yueting Zhuang, Lidong Bing

arXiv Stars197 🤗 Dataset downloads61,714 Total Pages 知乎 X

  • Interleaved image-text pretraining corpus from instructional videos
  • All the images and text are extracted from online instructional videos (22,000 class hours), covering multiple fundamental subjects, e.g., mathematics, physics, and chemistry.
  • Our textbook corpus providing a more coherent context and richer knowledge for image-text aligning.
  • More than 20,000 downloads within one month (Rank #2 in Huggingface Trending)

Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
Wenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Jiajun Liu, Weiming Lu, Peng Li, Yueting Zhuang

ACL 2026 arXiv Embodied-Omni Stars249 🤗 Dataset downloads17,708 Total Pages B站视频 机器之心 X 小红书

  • O1-style Embodied Reasoning Model
  • Interactive Embodied Scenario and Long-horizon Tasks
  • Autonomous Environment Exploration, Hidden Object Search, and Deep Reflection
  • Open-source dataset: 9.3K Observation-Reasoning-Action interleaved trajectories with 64K images
Outstanding Paper@ICLR LLM Agent workshop
sym

Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
Wenqi Zhang, Yongliang Shen, Weiming Lu, Yueting Zhuang

arXiv Stars1,512 Hugginface Spaces 知乎 机器之心

  • LLM-powered autonomous data analysis agent
  • Automated data querying, analysis, and visualization
  • Enterprise-level scenario
EMNLP 2024 Oral
sym

Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang

arXiv Stars85 Project 🤗 Dataset downloads13,358 Total 新智元 X AITime

  • Multimodal data engine
  • Synthetic massive abstract chart data
  • Enhance the abstract image perception and reasoning ability of multimodal models
ACL 2024
sym

Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
Wenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu

arXiv PaperWeekly

  • Investigate LLM’s self-reflection ability
  • Break the blind faith in LLM’s self-reflection ability
  • Inference time scale-up for better reasoning ability
ACL 2024
sym

Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, Weiming Lu

arXiv Stars134

  • Self-evolving LLM agent
  • Policy-level reflection and optimization
  • Dynamic environment and game scenarios
IJCAI 2022
sym

A Closed-Loop Perception, Decision-Making and Reasoning Mechanism for Human-Like Navigation
Wenqi Zhang, Kai Zhao, Peng Li, Xiao Zhu, Yongliang Shen, Yanna Ma, Yingfeng Chen, Weiming Lu

arXiv Stars46 YouTube Bilibili

  • Autonomous navigation framework for robots
  • Self-exploration for better navigation strategy
  • Action-to-State inverse reasoning process
arxiv2406
sym

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, Lidong Bing

arXiv Stars1,307 hf_space 🤗 Model downloads583,504 Total 🤗 Dataset downloads2,142 Total

  • Open-source Video-language model

🎖 Honors and Awards

  • 2025.06 Huawei TopMinds (华为天才少年)
  • 2024.10 National Scholarship (Top 1 %)
  • 2024.10 Distinguished Reviewer Award @33rd ACM International Conference on Information and Knowledge Management (CIKM 2024)
  • 2024.05 Outstanding Paper @ ICLR2024 LLM Agent Workshop (TOP 3%)
  • 2021-2022, 2022-2023, 2023-2024 Excellent Postgraduate Student Scholarship of Zhejiang University

💻 Experience

  • 2024.05, Research Intern, Alibaba DAMO Academy, Supervisor: Xin Li, Lidong Bing
    • Vision-language Pretraining
    • Developing video-language models with colleagues
  • 2020.02 - 2021.08, Algorithm Engineer, Advanced Institute of Information Technology, Peking University, Leader: Peng Li, Tao Wang
    • RL-based robot motion control and navigation framework

💬 Invited Talks

  • 2025.05, Embodied-Reasoner @智猩猩 [video]
  • 2024.10, Multimodal Self-instruct @AITime [video]
  • 2024.08, LLM Agent @MetaGPT Team (DeepWisdom)
  • 2024.08, How to Apply an LLM to Data Science@觅炽科技

📂 Services

Area Chair:

  • ACL, EMNLP (ACL-ARR) 2025

PC Member:

  • NIPS 2025, ICCV2025, IJCAI2025, ACL 2023-2025, ACM-MM 2024, ICLR 2024, WWW 2024, EMNLP 2023-2025, CIKM 2024