Dafeng Wei

Researcher · AgiBot (智元机器人)

I work on Vision-Language Models (VLM) and Vision-Language-Action models (VLA) for embodied AI — teaching robots to reason about the world and act in it.

Previously, I was a senior algorithm engineer on the autonomous driving team at Li Auto, where I was a core contributor to DriveVLM — an industry-first dual-system ("fast–slow") autonomous driving solution and the first of its kind deployed at scale on production vehicles — and built the multimodal retrieval systems powering the data loop. Before that I was an algorithm engineer at ByteDance. I received my M.S. from Shanghai Jiao Tong University, advised by Prof. Hongtao Lu.

Portrait of Dafeng Wei

News

Publications

* denotes equal contribution. Also see my Google Scholar profile.

2026

τ0-VLA framework: high-level policy with world-model-guided test-time computation driving a low-level VLA

τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Shanghai Innovation Institute & AgiBot (authors listed alphabetically, incl. Dafeng Wei)

Technical report, 2026

A hierarchical robot foundation model: the high-level planner performs world-model-guided beam search over subtasks with error-correcting execution memory, driving a cross-embodiment VLA — lifting success on up-to-12-minute real-robot tasks from 27.5% to 45.0%.

Core contributor — built the AgibotVLA training framework from scratch, ran the initial VLA pre-training, and developed the high-level VLM planner end-to-end: task formulation, data construction, training, and evaluation.

2025

GenieReasoner overview figure

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo

arXiv preprint, 2025

GenieReasoner jointly optimizes embodied reasoning and action execution, with the ERIQ reasoning benchmark (6k+ QA pairs) and FACT, a flow-matching action tokenizer.

AgiBot World Colosseo overview figure

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

AgiBot-World-Contributors (incl. Dafeng Wei)

IROS 2025 🏆 Best Paper Award Finalist · IEEE T-RO 2026

An open platform with 1M+ real-robot trajectories across 217 tasks, and GO-1, a generalist policy built on latent action representations.

BEV-TSR overview figure

BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving

Tao Tang*, Dafeng Wei*, Zhengyu Jia, Tian Gao, Changwei Cai, Chengkai Hou, Peng Jia, Kun Zhan, Haiyang Sun, Jingchen Fan, Yixing Zhao, Fu Liu, Xiaodan Liang, Xianpeng Lang, Yang Wang

AAAI 2025

Retrieving complex driving scenes with free-form text queries directly in BEV feature space.

2022

Efficient SlowFast thumbnail

Efficient Dual Attention SlowFast Networks for Video Action Recognition

Dafeng Wei, Ye Tian, Liqing Wei, Hong Zhong, Siqian Chen, Shiliang Pu, Hongtao Lu

Computer Vision and Image Understanding (CVIU), 2022

Lightweight two-stream video networks with a cross-modality dual attention fusion module (CMDA).

2021

Active learning pipeline figure

Towards Dynamic and Scalable Active Learning with Neural Architecture Adaption for Object Detection

Fuhui Tang*, Dafeng Wei*, Chenhan Jiang, Hang Xu, Andi Zhang, Wei Zhang, Hongtao Lu, Chunjing Xu

BMVC 2021

Active learning for object detection that adapts network architecture as the labeled pool grows; deployed on large-scale autonomous-driving data at Huawei.

2020

TableCell dataset overview from the ICPR poster

Image-based Table Cell Detection: a Novel Table Structure Decomposition Method with New Dataset

Dafeng Wei, Hongtao Lu, Yi Zhou, Kai Chen

ICPR 2020

Detecting table cells as objects to recover table structure, with TableCell — an open dataset of 170K cell-level annotations.

Experience

Internships

Education

Miscellaneous