兰州理工大学学报 ›› 2026, Vol. 52 ›› Issue (3): 91-101.

• 自动化技术与计算机技术 • 上一篇    下一篇

基于深度强化学习的交通信号优先控制

王志文*1,2,3, 于宇凌1, 杨康康1, 王浩旭1, 苗葳4   

  1. 1.兰州理工大学 自动化与电气工程学院, 甘肃 兰州 730050;
    2.兰州理工大学 甘肃省工业过程先进控制重点实验室, 甘肃 兰州 730050;
    3.兰州理工大学 电气与控制工程国家级实验教学示范中心, 甘肃 兰州 730050;
    4.甘肃紫光智能交通与控制技术有限公司, 甘肃 兰州 730030
  • 收稿日期:2024-02-28 出版日期:2026-06-28 发布日期:2026-06-30
  • 通讯作者: 王志文(1976-),男,甘肃武威人,博士,教授,博导.Email:wzw@lut.edu.cn
  • 基金资助:
    国家自然科学基金(62263019),甘肃省青年科技基金计划(21JR7A346),甘肃省重大专项(21ZD4GA028)

Traffic signal priority control based on deep reinforcement learning

WANG Zhi-wen1,2,3, YU Yu-ling1, YANG Kang-kang1, WANG Hao-xu1, MIAO Wei4   

  1. 1. School of Automation and Electrical Engineering, Lanzhou University of Technology, Lanzhou 730050, China;
    2. Key Laboratory of Gansu Advanced Control for Industrial Processes, Lanzhou University of Technology, Lanzhou 730050, China;
    3. National Demonstration Center for Experimental Electrical and Control Engineering Education, Lanzhou University of Technology, Lanzhou 730050, China;
    4. Gansu Ziguang Intelligent Transportation and Control Technology Co., Ltd., Lanzhou 730030, China
  • Received:2024-02-28 Online:2026-06-28 Published:2026-06-30

摘要: 现有的交通信号优先控制系统无法有效适应日益复杂多变的城市交通环境,并且在实施交通信号优先控制时,容易对非优先车辆造成不必要的延误.针对这种不足,从控制策略和优化目标两个方面着手,基于交叉路口的实时交通信息,并结合深度强化学习高效处理高维连续数据和自学习的特点,提出一种基于深度强化学习的优先控制策略(G-D3QN),改进基于竞争架构的深度双Q网络对信号优先控制策略进行求解.同时,构建双目标优化奖励函数,兼顾公交车辆和社会车辆的通行需求,在最大化交叉路口通行能力的同时,降低公交车辆延误.最后,基于SUMO搭建仿真环境,将改进后的算法与D3QN、DDQN、Q-Learning三种基类算法进行对比,结果表明改进算法能够有效协同提高公交车和社会车辆的通行效率.

关键词: 交通信号优先, 交通信号控制, 深度强化学习, SUMO仿真平台

Abstract: The existing traffic signal priority control system cannot effectively adapt to the increasingly complex and changing urban traffic environment, and is prone to cause unnecessary delays to non-priority vehicles when implementing traffic signal priority control. To address this shortcoming, this study proposes improvements from both the control strategy and optimization objectives. Based on the real-time traffic information of intersections and combining the characteristics of deep reinforcement learning to efficiently process high-dimensional continuous data and self-learning, a priority control strategy based on deep reinforcement learning is proposed to improve the deep double-Q network based on a dueling architecture to solve the signal priority control strategy. At the same time, a dual-objective optimization incentive function is constructed to balance the access needs of both public transport vehicles and general traffic,aiming to reduce delays of public transport vehicles while maximizing the capacity of intersections. Finally, a simulation environment is built using the SUMO to compare the improved algorithm with the three base class algorithms of D3QN,DDQN and Q-Learning. The experimental results show that the improved algorithm can effectively improve the access efficiency of buses and social vehicles in a coordinated manner.

Key words: transit signal priority, traffic signal control, deep reinforcement learning, SUMO simulation platform

中图分类号: