👋 Hi there! I am Zhiwei Shang, currently a second-year Computer Science Ph.D. student in the School of Data Science at The Chinese University of Hong Kong, Shenzhen, supervised by Prof. Zhongxiang Dai. I also work as an LLM Agent Researcher at DeepWisdom.
Previously, I was a research assistant worked with Prof. Meixin Zhu at The Hong Kong University of Science and Technology (Guangzhou) and Dr. Chenjia Bai at the Shanghai Artificial Intelligence Laboratory. I received my M.E. in Computer Technology from the University of Chinese Academy of Sciences in 2023, advised by Prof. Yunduan Cui, and my B.E. in Hydrology and Water Resources Engineering from Sichuan University in 2020.
My current research interests mainly lie in Large Language Models (LLMs), LLM-based Agents and Reinforcement Learning (RL). I believe that the development of Artificial Intelligence can help us build a better, safer and more equal society!You’re more than welcome to write me an email to connect (or make friends)! Let’s explore the exciting world together! 🌌
📖 Education
-
Sep. 2025 - present
Ph.D. student in Computer Science
The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)
Advisor: Prof. Zhongxiang Dai
Research: Large language models and reinforcement learning -
Sep. 2020 - Jun. 2023
M.E. in Computer Technology
University of Chinese Academy of Sciences (UCAS)
Advisor: Prof. Yunduan Cui
Research: Reinforcement learning theory and applications
Thesis: Reinforcement Learning Control Methods Based on Relative Entropy Regularization -
Sep. 2016 - Jun. 2020
B.E. in Hydrology and Water Resources Engineering
Sichuan University (SCU)
Thesis: Deep Learning-Based Runoff Prediction in the Upper Minjiang River Basin
💼 Research Experience
-
2026.07 - present, LLM Agent Researcher @DeepWisdom, advised by Chenglin Wu.
-
2023.08 - 2025.06, Research Assistant @Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Meixin Zhu.
-
2024.02 - 2024.09, Research Assistant @Shanghai Artificial Intelligence Laboratory, advised by Dr. Chenjia Bai.
-
2020.09 - 2023.06, Core Member & Research Assistant @Shenzhen Institute of Advanced Technology, advised by Prof. Yunduan Cui.
🔥 News
- 2026.07: I joined DeepWisdom as an LLM Agent Researcher.
- 2026.06: 🎉 Three papers were accepted to ICML 2026 workshops.
- 2026.05: 🎉 Received the ICML 2026 Silver Reviewer Award.
- 2026.04: 🎉 Three papers were accepted to ICML 2026, including one Spotlight (Top 2.2%).
- 2025.09: 👨🎓 I was admitted to the Ph.D. program at the School of Data Science, The Chinese University of Hong Kong, Shenzhen.
- 2025.06: 🎉 Our paper was accepted to IROS 2025.
- 2025.05: 🎉 Our paper was accepted by IEEE Transactions on Intelligent Transportation Systems (TITS).
- 2024.03: 🎉 Our paper was accepted to IV 2024.
📝 Publications & Patents
* Co-first author, ✉️ Corresponding author.
Working Papers
-
Effective Reinforcement Learning Control using Conservative Soft Actor-Critic.
Zhiwei Shang* , Xinyi Yuan*, Wenjun Huang, Yunduan Cui, Di Chen, Meixin Zhu✉️.
Preprint, 2025 -
FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits.
Pingchen Lu, Zhi Hong, Zhiwei Shang , Zhiyong Wang, Yikun Ban, Yao Shu, Min Zhang, Shuang Qiu, Zhongxiang Dai✉️.
Preprint, 2025 -
A Rugged Compass in Flatland: Curvature-Aware Rectification for Secure Test-Time Adaptation.
Mingrong Gong, Junhao Dong, Zhiwei Shang , Daizong Liu, Siheng Wang, Zhengtao Yao, Ting Peng, Sergio Escalera, Xinghua Qu, Yew-Soon Ong✉️.
Preprint, 2026 -
VLM safety via geometry reference.
Mingrong Gong, Junhao Dong, Xuanhui Lin, Zhiwei Shang , Jiaming Zhang, Daizong Liu, Yikai Wang, Ting Peng, Sergio Escalera, Xinghua Qu, Yew-Soon Ong✉️.
Preprint, 2026
Published / Accepted Papers
-
Fusion is the New Mutation: Bandit-Guided Evolution on Workflow Graphs.
Zhiwei Shang , Jiahang Sun, Mingrong Gong, Mingze Kong, Qu Zikun, Pingchen Lu, Junhao Dong, Zhipiao Liu, Hongwei Yang, Guoqing Xie, Yao Shu, Zhongxiang Dai✉️.
ICML 2026 Workshop on Compositional Learning: Safety, Interpretability, and Agents. -
CB-Orchestrator: Adaptive Workflow Optimization for LLM Agents via Contextual Bandits.
Jiahang Sun, Zhiwei Shang , Zhipiao Liu, Hongwei Yang, Guoqing Xie, Shuang Qiu, Zhongxiang Dai✉️.
ICML 2026 Workshop on Compositional Learning: Safety, Interpretability, and Agents. -
Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction.
Mingze Kong, Zikun Qu, Zhongquan Zhou, Pengyu Liang, Xiang Li, Zhiwei Shang , Zhi Hong, Kaiyu Huang, Zhiyong Wang, Zhongxiang Dai✉️.
ICML 2026 Workshop on RL from World Feedback. -
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks.
Zhi Hong*, Qian Zhang*, Jiahang Sun, Zhiwei Shang, Mingze Kong, Xiangyi Wang, Yao Shu, Zhongxiang Dai✉️.
ICML 2026 (Spotlight, Top 2.2%) -
T-POP: Test-Time Personalization with Online Preference Feedback.
Qu Zikun, Min Zhang, Mingze Kong, Xiang Li, Zhiwei Shang, Zhiyong Wang, Yikun Ban, Shuang Qiu, Yao Shu, Zhongxiang Dai✉️.
ICML 2026 -
Social Hippocampus Memory Learning.
Liping Yi, Zhiming Zhao, Kewen Zhu, Xiang Li, Zhiwei Shang, Qinghua Hu✉️.
ICML 2026 -
Preference Aligned Diffusion Planner for Quadrupedal Locomotion Control.
Xinyi Yuan*, Zhiwei Shang* , Zifan Wang, Chenkai Wang, Zhao Shan, Meixin Zhu✉️, Chenjia Bai✉️, Weiwei Wan, Kensuke Harada, Xuelong Li.
IROS 2025, Oral Presentation | [Website] -
Dynamic High-Order Control Barrier Functions with Diffuser for Safety-Critical Trajectory Planning at Signal-Free Intersections.
Di Chen, Ruiguo Zhong, Kehua Chen, Zhiwei Shang, Meixin Zhu✉️, Edward Chung.
IEEE Transactions on Intelligent Transportation Systems(TITS), 2025 -
Learning Realistic and Reactive Traffic Agents.
Meixin Zhu✉️, Di Chen, Xinyi Yuan, Zhiwei Shang, Chenxi Liu.
IV 2024 -
Relative Entropy Regularized Sample-Efficient Reinforcement Learning with Continuous Actions.
Zhiwei Shang, Renxing Li, Chunhua Zheng, Huiyun Li, Yunduan Cui✉️.
IEEE Transactions on Neural Networks and Learning Systems(TNNLS), 2023 -
Efficient distributional reinforcement learning with Kullback-Leibler divergence regularization.
Renxing Li, Zhiwei Shang, Chunhua Zheng, Huiyun Li, Qing Liang, Yunduan Cui✉️.
Applied Intelligence, 2023 -
Dynamic Policy Programming with Descending Regularization for Efficient Reinforcement Learning Control.
Renxing Li, Zhiwei Shang, Chunhua Zheng, Huiyun Li, Qing Liang, Yunduan Cui✉️.
PRAI 2022 -
Shiftable Dynamic Policy Programming for Efficient and Robust Reinforcement Learning Control.
Zhiwei Shang, Huiyun Li, Yunduan Cui✉️.
ROBIO 2021
Patents
- Zhiwei Shang , Yunduan Cui, Zhengkun Yi, Xiang Xie, Huiyun Li. AC framework based on relative entropy regularization and its application to control robotic arms. Chinese invention patent application, publication no. CN116128017A.
🎖 Awards
- May 2026, ICML 2026 Silver Reviewer Award.
- Dec. 2022, Director’s Innovation Award (Outstanding Graduate Student Award), Shenzhen Institute of Advanced Integration Technology, Chinese Academy of Sciences and The Chinese University of Hong Kong (17 recipients across the institute).
- May 2019, Outstanding Project Award, Sichuan Provincial College Student Innovation and Entrepreneurship Program (Top 10%).
💬 Talks
- Apr. 2024, Guest Lecture on Reinforcement Learning, The Hong Kong University of Science and Technology (Guangzhou).
💻 Services
- Program Committee Member: ICML 2026 Workshop on From Frames to Stories (F2S)
- Conference Reviewer: ICML, NeurIPS, ICLR, AAAI, EMNLP, ICRA, IROS
- Journal Reviewer: IEEE Transactions on Intelligent Vehicles (TIV), Applied Intelligence
- Teaching Assistant, School of Data Science, The Chinese University of Hong Kong, Shenzhen:
- CSC3100 Data Structures (Summer 2026)
- DDA 2001 Introduction to Data Science (Spring 2026)
- CSC 4303 Network Programming (Spring 2026)