走进学院
学院资讯
科研动态
学生培养
邓悦
邓悦

北京中关村学院导师

邓悦,浙江大学博士。主要从事强化学习、多智能体强化学习、决策大模型等方面研究工作。投稿发表国际顶会顶刊论文JMLR/IJCAI/NeurIPS/TOSEM等10余篇,参与国家自然科学基金面上项目“面向多模态知识搜索的多轮机器阅读理解和持续学习研究”和浙江省自然科学基金重点项目“低资源跨语言多模态内容表示学习技术研究”项目。现研究基于大语言模型的多智能体合作以及大规模智能体协同进化与组织演进等题目。

个人主页链接:https://devindeng94.github.io/

研究方向

1. 强化学习

2. 多智能体强化学习

3. 决策大模型

 

教育经历

2020-2025 浙江大学计算机技术专业 博士

2018-2020 墨尔本大学信息技术专业 硕士

2013-2017大连理工大学自动化专业 学士

 

工作经历

2025-至今 中关村人工智能研究院 研究员

 

代表性学术论文

Deng, Y., Wang, Z., Chen, X., & Zhang, Y. (2023). Boosting Multi-agent Reinforcement Learning via Contextual Prompting. Journal of Machine Learning Research24(399), 1-34.

Deng, Y., Wang, Z. R., & Zhang, Y. (2024, August). Improving multi-agent reinforcement learning with stable prefix policy. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence. Jeju, South Korea: IJCAI (pp. 49-57).

Wang, Z., Deng, Y., Long, J., & Zhang, Y. (2024). Parallelizing model-based reinforcement learning over the sequence length. Advances in Neural Information Processing Systems37, 131398-131433.

Lin, Z., Li, G., Zhang, J., Deng, Y., Zeng, X., Zhang, Y., & Wan, Y. (2022). Xcode: Towards cross-language code representation with large-scale pre-training. ACM Transactions on Software Engineering and Methodology (TOSEM)31(3), 1-44.

Deng, Y., Ma, W., Fan, Y., Song, R., Zhang, Y., Zhang, H., & Zhao, J. (2024). Smac-r1: The emergence of intelligence in decision-making tasks. arXiv preprint arXiv:2410.16024.

上一篇    邓岳
董彬    下一篇