北京中关村学院导师
个人主页链接:https://devindeng94.github.io/
研究方向
1. 强化学习
2. 多智能体强化学习
3. 决策大模型
教育经历
2020-2025 浙江大学计算机技术专业 博士
2018-2020 墨尔本大学信息技术专业 硕士
2013-2017大连理工大学自动化专业 学士
工作经历
2025-至今 中关村人工智能研究院 研究员
代表性学术论文
Deng, Y., Wang, Z., Chen, X., & Zhang, Y. (2023). Boosting Multi-agent Reinforcement Learning via Contextual Prompting. Journal of Machine Learning Research, 24(399), 1-34.
Deng, Y., Wang, Z. R., & Zhang, Y. (2024, August). Improving multi-agent reinforcement learning with stable prefix policy. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence. Jeju, South Korea: IJCAI (pp. 49-57).
Wang, Z., Deng, Y., Long, J., & Zhang, Y. (2024). Parallelizing model-based reinforcement learning over the sequence length. Advances in Neural Information Processing Systems, 37, 131398-131433.
Lin, Z., Li, G., Zhang, J., Deng, Y., Zeng, X., Zhang, Y., & Wan, Y. (2022). Xcode: Towards cross-language code representation with large-scale pre-training. ACM Transactions on Software Engineering and Methodology (TOSEM), 31(3), 1-44.
Deng, Y., Ma, W., Fan, Y., Song, R., Zhang, Y., Zhang, H., & Zhao, J. (2024). Smac-r1: The emergence of intelligence in decision-making tasks. arXiv preprint arXiv:2410.16024.



