I’m Zhen Huang (“黄臻” in Chinese), a second-year Ph.D. student at Fudan University, under the supervision of Pengfei Liu (affiliated with GAIR Lab @ SJTU). My current research interest is long-horizon agents. Previously, I also focused on data-centric AI, in particular curating high-quality data in a more automated way across various stages of pre-training, mid-training and post-training. I’m currently an intern at the
Hunyuan LLM Agent Team, Tencent.
🔥 News
-
2026.04: 🎉🎉 One paper accepted by ACL 2026 (main) on LLM Pretraining - “Unlocking the Value of Scientific Data for Pre-training”.
-
2025.07: 🎉🎉 One paper accepted by COLM 2025 on LLM reasoning - “LIMO: Less is More for Reasoning”.
-
2024.09: 🎉🎉 One paper accepted by NeurIPS 2024 on LLM&LMM evaluation - “OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI”.
📝 Publications
( * : equal contribution, † : corresponding author )
-
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
Zhen Huang, Yikun Wang, Shijie Xia, Pengfei Liu†
arXiv 2026
Summary: A learned orchestrator tailors data curation to each pretraining example, improving data quality with less processing compute.
-
SciPedia: Unlocking the Value of Scientific Data for Pre-training
Yiwei Qin*, Zhen Huang*, Tiantian Mi*, Weiye Si, Qipeng Guo, Siyuan Feng, Pengfei Liu†
ACL 2026 (main)
[Paper][Code][🤗 Dataset][🤗 Model (3B)][🤗 Model (7B)][🤗 Benchmark]
Summary: SciPedia is a 900B-token scientific corpus enriched through cleaning and pedagogical augmentation for more effective pretraining.
-
LIMO: Less is More for Reasoning
Yixin Ye*, Zhen Huang*, Yang Xiao, Ethan Chern, Shijie Xia, Pengfei Liu†
COLM 2025
[Paper][Code][Model][🤗 Datasets][Featured by AK][机器之心][Talk (Chinese)]
Summary: Strong mathematical reasoning can emerge from supervised fine-tuning on a small set of carefully curated examples.
-
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Zhen Huang, Zengzhi Wang, Shijie Xia, Xuefeng Li, Haoyang Zou, Ruijie Xu, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, Yikai Zhang, Yuqing Yang, Ting Wu, Binjie Wang, Shichao Sun, Yang Xiao, Yiyuan Li, Fan Zhou, Steffi Chern, Yiwei Qin, Yan Ma, Jiadi Su, Yixiu Liu, Yuxiang Zheng, Shaoting Zhang†, Dahua Lin†, Yu Qiao†, Pengfei Liu†
NeurIPS 2024
[Paper][Code][Homepage][🤗 Datasets][ 🤗 Competition][Featured by AK][机器之心][量子位]
Summary: OlympicArena is a multimodal reasoning benchmark with 11,163 bilingual Olympiad problems across seven disciplines.
( * : equal contribution, † : corresponding author )
-
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
Zhen Huang, Yikun Wang, Shijie Xia, Pengfei Liu†
arXiv 2026
Summary: A learned orchestrator tailors data curation to each pretraining example, improving data quality with less processing compute.
-
AdaCodec: A Predictive Visual Code for Video MLLMs
Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si, Chenglin Li, Shuai Dong, Kele Shao, Ruilin Li, Dianyi Wang, Nan Duan, Jiaqi Wang†
NeurIPS 2026 (spotlight)
Summary: Compact representations of inter-frame changes reduce redundant visual tokens and improve video understanding.
-
DaVinci-Dev: Agent-native Mid-training for Software Engineering
Ji Zeng, Dayuan Fu, Tiantian Mi, Yumin Zhuang, Yaxing Huang, Xuefeng Li, Lyumanshan Ye, Muhang Xie, Qishuo Hua, Zhen Huang, Mohan Jiang, Hanning Wang, Jifan Lin, Yang Xiao, Jie Sun, Yunze Wu, Pengfei Liu†
ICML 2026 (Oral)
[Paper][Code][🤗 Dataset][🤗 Model (32B)][🤗 Model (72B)]
Summary: Agent-native mid-training uses realistic development trajectories to strengthen software agents’ repository-level problem solving.
-
SciPedia: Unlocking the Value of Scientific Data for Pre-training
Yiwei Qin*, Zhen Huang*, Tiantian Mi*, Weiye Si, Qipeng Guo, Siyuan Feng, Pengfei Liu†
ACL 2026 (main)
[Paper][Code][🤗 Dataset][🤗 Model (3B)][🤗 Model (7B)][🤗 Benchmark]
Summary: SciPedia is a 900B-token scientific corpus enriched through cleaning and pedagogical augmentation for more effective pretraining.
-
InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative LLM Research
Yunze Wu, Dayuan Fu, Weiye Si, Zhen Huang, Mohan Jiang, Keyu Li, Shijie Xia, Jie Sun, Tianze Xu, Xiangkun Hu, Pengrui Lu, Xiaojie Cai, Lyumanshan Ye, Wenhong Zhu, Yang Xiao, Pengfei Liu†
ICLR 2026
[Paper][Code][🤗 Dataset][OpenReview]
Summary: InnovatorBench evaluates the full LLM research workflow across 20 tasks in an executable research environment.
-
LIMO: Less is More for Reasoning
Yixin Ye*, Zhen Huang*, Yang Xiao, Ethan Chern, Shijie Xia, Pengfei Liu†
COLM 2025
[Paper][Code][Model][🤗 Datasets][Featured by AK][机器之心][Talk (Chinese)]
Summary: Strong mathematical reasoning can emerge from supervised fine-tuning on a small set of carefully curated examples.
-
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Zhen Huang, Zengzhi Wang, Shijie Xia, Xuefeng Li, Haoyang Zou, Ruijie Xu, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, Yikai Zhang, Yuqing Yang, Ting Wu, Binjie Wang, Shichao Sun, Yang Xiao, Yiyuan Li, Fan Zhou, Steffi Chern, Yiwei Qin, Yan Ma, Jiadi Su, Yixiu Liu, Yuxiang Zheng, Shaoting Zhang†, Dahua Lin†, Yu Qiao†, Pengfei Liu†
NeurIPS 2024
[Paper][Code][Homepage][🤗 Datasets][ 🤗 Competition][Featured by AK][机器之心][量子位]
Summary: OlympicArena is a multimodal reasoning benchmark with 11,163 bilingual Olympiad problems across seven disciplines.
📖 Educations
- 2025.09 - 2030.06 (Expected), Ph.D. in Computer Science and Technology, Fudan University, Shanghai, China
- Research focus: Large Language Models
- 2021.09 - 2025.07, B.Eng. in Software Engineering, Soochow University, Suzhou, China
- GPA: 4.0/4.0, Ranking: 1/95
🌐 Services
- Reviewer: NeurIPS 2025, ICLR 2026, ICML 2026
🥥 Misc
- ⚽️ I’m a passionate football enthusiast. I’m a devoted Chelsea FC supporter, drawn to the club’s legendary spirit—a testament to resilience, determination, and never giving up. Beyond the club itself, I’m an ardent admirer of José Mourinho, whose distinctive personality and tactical genius (counter-attacking) have left an indelible mark on football.
AlphaXiv