CORTEXA
← Browse
arxiveess.SY2026-07-10

A Distributionally Robust Multi-agent Reinforcement Learning Framework for Intelligent Intersection Control

Shuwei Pei, Joran Borger, Arda Kosay, Bayu Jayawardhana, Muhammed O. Sayin, Saeed Ahmed

Multi-agent reinforcement learning (MARL) has emerged as a promising approach for traffic signal control. However, standard MARL policies typically optimize for expected returns under nominal conditions, leaving them highly vulnerable to spatial-temporal demand shifts and catastrophic congestion under adverse scenarios. To address this critical limitation, this paper proposes an algorithm-agnostic Distributionally Robust (DR) MARL framework integrating an adaptive Contextual-Bandit Worst-Case Estimator (CB-WCE). Operating on a slower timescale, the CB-WCE co-evolves with the traffic controllers by dynamically generating adversarial demand mixtures during training. This steers the learning process to fortify policies against bottleneck scenarios without requiring modifications to the underlying MARL architectures. The framework is evaluated across value-based, actor-critic, and policy-gradient methods on both a synthetic 5x5 grid and a heterogeneous Monaco City network. Empirical results demonstrate that the DR framework prevents unbounded queue growth and profoundly enhances both worst-case robustness and average-case efficiency. Notably, for the Proximal Policy Optimization (PPO) architecture in the Monaco environment, on average, robust retraining reduced the worst-case queue length by 74.39% and improved the average-case network-wide queue length by 75.45%. Furthermore, the retrained policies exhibit strong zero-shot generalization to unseen traffic distributions, highlighting the framework's scalability and potential for resilient real-world urban deployment.

View free PDFSource page

Related papers

arxiveess.SY2026-07-17Cited by 9

A PI+R Control Scheme Based on Multi-agent Systems for Economic Dispatch in Isolated BESSs

Yalin Zhang, Zhongxin Liu, Zengqiang Chen

Battery energy storage systems (BESSs) are widely used in smart grids. However, power consumed by inner impedance and the capacity degradation of each battery unit become particularly severe, which has resulted in an increase in operating costs. The general economic dispatch (ED)…

View free PDFSource page
arxivcs.MAeess.SYmath.OC2026-07-18

Laplacian Spectral Shaping for Non-Uniform Scaling Formation Control of Open Multi-Agent Systems

Tao He, Gangshan Jing

Non-uniform scaling control enables a multi-agent formation to adjust its shape by compressing or stretching independently along different coordinate axes through inter-agent interactions, offering high flexibility in complex environments. The fundamental idea is encoding the des…

View free PDFSource page
arxiveess.SY2026-07-20

On Optimal Event-Triggered Distributed Control for Stochastic Multi-Agent Systems via Reinforcement Learning

Ziming Wang, Bingbing Li, Karl H. Johansson, Apostolos I. Rikos

We propose a reinforcement learning (RL) based optimal distributed control algorithm for the multi-agent systems (MASs) with stochastic uncertainties. Unlike existing methods, during the optimized backstepping design process, we use the actor-critic-identifier structure. The acto…

View free PDFSource page
arxivcs.MAcs.AIcs.CYcs.DCeess.SY2026-07-19

The Optimization Trilemma: Efficiency, Comfort and Fairness in Decentralized Multi-agent Coordination

Jovan Nikolic, Maciej Krzysztof Zuziak, Evangelos Pournaras

The problem of fair multi-agent coordination in decentralized settings is one of the most pressing challenges for building efficient collaborative systems. Resource allocation is based on optimized collective arrangements accounting for agents' needs. Such coordination should not…

View free PDFSource page
arxiveess.SYcs.MAmath.OC2026-07-21

How network perturbations distort agreement trajectories in LTI multi-agent systems

Gal Barkai, Irinel-Constantin Morărescu

Distributed coordination of multi-agent systems frequently relies on cooperative protocols designed to achieve agreement on a prescribed, non-trivial trajectory. While the robustness of such protocols to various uncertainties is well documented, existing literature universally as…

View free PDFSource page