I am a master’s student in Finance at the Guanghua School of Management, Peking University, and received my B.S. in Artificial Intelligence from the School of Data Science, Fudan University, in 2025 (ranked 2nd in my major at graduation). My research mainly focuses on LLM agents (harness engineering) and post-training. I am currently an Algorithm Researcher at NEX-AGI (Shanghai Qiji Zhifeng Co., Ltd.), working on the self-evolution of LLMs and multi-agent systems. Before that, I spent several years in quantitative finance, working on CTA / index-futures strategies and high-frequency factor mining — which is exactly why my master’s degree is in Finance.我是北京大学光华管理学院金融专业的硕士研究生,并于 2025 年在复旦大学大数据学院获得人工智能学士学位(毕业时专业排名第二)。我的研究方向主要为大语言模型智能体(Harness Engineering)与后训练。我目前在 NEX-AGI(上海奇绩智峰)担任算法研究员,进行大语言模型的自演化与多智能体相关研究。在此之前,我曾在量化金融领域工作数年,从事 CTA / 股指期货策略与高频因子挖掘——这就是为什么我硕士专业为金融。
You can find my publications on Google Scholar. Feel free to reach out to me at cjpan25@stu.pku.edu.cn.你可以在 Google Scholar 上查看我的论文。欢迎通过邮箱 cjpan25@stu.pku.edu.cn 与我联系。
🔥 News最新动态
- 2026.06: 🏆 AHE ranks #1 on the Terminal-Bench 2.0 leaderboard!🏆 AHE 登上 Terminal-Bench 2.0 排行榜 榜首!
- 2026.04: 🎉 Agentic Harness Engineering (AHE) is released on arXiv.🎉 Agentic Harness Engineering (AHE) 已在 arXiv 发布。
- 2026.04: 🎉 EVPO is released on arXiv.🎉 EVPO 已在 arXiv 发布。
- 2026.04: 🎉 AutoJudger is accepted by ACL 2026 (Main Conference)!🎉 AutoJudger 被 ACL 2026 主会 录用!
- 2026.02: 🎉 AgentCPM-Explore is released on arXiv.🎉 AgentCPM-Explore 已在 arXiv 发布。
💻 Internships实习经历
2026.1 - Present · LLM Algorithm Research, NEX-AGI (Shanghai Qiji Zhifeng Co., Ltd.), Shanghai2026.1 - 至今 · 大模型算法研究,NEX-AGI(上海奇绩智峰),上海
Self-evolution of LLMs and multi-agent systems.大语言模型的自演化与多智能体方向。
2025.10 - 2026.01 · LLM Algorithm Research, THUNLP, Tsinghua University, Beijing2025.10 - 2026.01 · 大模型算法研究,清华大学 THUNLP,北京
AgentCPM-Explore model development; improving RL algorithms for LLMs in multi-turn tool-use settings.AgentCPM-Explore模型研发,面向多轮工具学习场景的大模型强化学习算法改进。2025.03 - 2025.07 · Quantitative Strategy Research, GenWealth Capital, Shanghai2025.03 - 2025.07 · 量化策略研究,华钧广汇,上海
Intraday quantitative strategies for index futures and options, and high-frequency factor mining.股指期货期权日内量化策略与高频因子挖掘。
2024.07 - 2024.11 · Quantitative Strategy Research, Zhanhong Investment, Shanghai2024.07 - 2024.11 · 量化策略研究,展弘投资,上海
Intraday quantitative trading strategies for commodity and index futures.商品期货/股指期货的日内量化交易策略。
📝 Publications论文发表

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
Jiahang Lin*, Shichun Liu*, Chengjun Pan*, Lizhi Lin, Shihan Dou, Xuanjing Huang, Hang Yan, Zhenhua Han†, Tao Gui†
- AHE is an observability stack for the automatic optimization of coding-agent harnesses, with three pillars: component observability (NexAU), experience observability (Agent Debugger), and decision observability (evidence-driven Evolve Agent).AHE 是一套用于自动优化编码智能体 harness 的可观测性体系,包含三大支柱:组件可观测性(NexAU)、经验可观测性(Agent Debugger)以及决策可观测性(基于证据的 Evolve Agent)。
- Without changing the model, AHE pushes Terminal-bench 2 from 69.7% to 77.0% across iterations, with strong cross-task and cross-model generalization.在不改动模型的前提下,AHE 通过多轮迭代将 Terminal-bench 2 从 69.7% 提升至 77.0%,并展现出强的跨任务、跨模型泛化能力。
|
|

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
Chengjun Pan*, Shichun Liu*, Jiahang Lin*, Dingwei Zhu, Jiazheng Zhang, Shihan Dou, Songyang Gao, Zhenhua Han, Binghai Wang, Rui Zheng, Xuanjing Huang†, Tao Gui†, Yansong Feng†
- We cast baseline selection in LLM post-training as a Kalman filtering problem, unifying PPO and GRPO as two extremes of the Kalman gain, and prove that the sign of explained variance (EV) is the exact boundary separating the variance-reducing from the variance-inflating critic regime.我们将大模型后训练中的 baseline 选择建模为卡尔曼滤波问题,把 PPO 与 GRPO 统一为卡尔曼增益的两个极端,并证明 explained variance(EV)的符号正是区分「降方差」与「增方差」critic 区间的精确边界。
- EVPO adaptively switches between critic-based and batch-mean advantage estimation per step based on EV sign, achieving the best results across Sokoban, FrozenLake, WebShop, and MATH.EVPO 依据每一步的 EV 符号,在基于 critic 与基于 batch 均值的优势估计之间自适应切换,在 Sokoban、FrozenLake、WebShop 与 MATH 上均取得最佳结果。
|
|

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
Xuanwen Ding*, Chengjun Pan*, Zejun Li, Jiwen Zhang, Siyuan Wang, Zhongyu Wei
- An agent-driven framework that uses Item Response Theory (IRT) to model item difficulty and model ability, adaptively selecting a small subset of benchmark items to match full-benchmark evaluation of MLLMs at a fraction of the cost.一个由智能体驱动的框架,利用项目反应理论(IRT)建模试题难度与模型能力,动态选取少量评测题目,即可以极低成本达到媲美全量评测的多模态大模型评估效果。
|

AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents
Heyang Chen, Xin Cong, …, Chengjun Pan, et al.

Exploring Systemic Risk Dynamics in the Chinese Stock Market: A Network Analysis with Risk Transmission Index
Xiaowei Zeng, Yifan Hu, Chengjun Pan, Yanxi Hou
🎖 Honors and Awards荣誉奖项
- 2025, Shanghai Outstanding Graduate.上海市优秀毕业生。
- 2024, 1st Prize, 2024 Tencent AI Arena Global Open Competition · Agent Game Algorithm Track · Mainland China Regional Final. Team: 五角场三分王, Fudan University.一等奖,2024 腾讯 AI Arena 全球公开赛 · 智能体博弈算法赛道 · 中国大陆赛区总决赛。队伍:五角场三分王,复旦大学。
- 2023, National Scholarship (the only recipient in the school that year).国家奖学金(学院该年唯一)。
- 2023, National Second Prize (top 1.5%), China Undergraduate Mathematical Contest in Modeling (CUMCM).全国大学生数学建模竞赛全国二等奖(前1.5%)。
- 2023, Shanghai First Prize, National Statistical Modeling Competition.全国大学生统计建模竞赛上海赛区一等奖。
- 2022 – 2024, Fudan University Outstanding Student.复旦大学优秀学生。
📖 Education教育经历
- 2025.09 - Present, M.S. in Finance, Guanghua School of Management, Peking University.金融硕士,北京大学光华管理学院。
- 2021.09 - 2025.06, B.S. in Artificial Intelligence, School of Data Science, Fudan University.人工智能学士,复旦大学大数据学院。