CORTEXA
← Browse
arxiveess.SP2026-07-10

Fused Constrained Policy Reuse Optimization for Wireless Resource Allocation

An Liu, Zheyuan Zhou, Kexuan Wang

Deep reinforcement learning (DRL) has been widely adopted for wireless resource allocation due to its model-free adaptability. However, online exploration is costly, as randomly initialized policies may violate long-term constraints before sufficient data are collected. Future wireless systems must cope with increasingly dynamic traffic, fluctuating channel conditions, and stringent energy efficiency requirements, demanding algorithms that can learn quickly with minimal environment interactions to reduce both energy consumption and signaling overhead. We develop Fused-CPRO, a knowledge-fused constrained policy reuse optimization method addressing these challenges. Fused-CPRO constructs the allocation policy as a mixture of a learnable target policy, source policies from related scenarios, and domain-knowledge (DK) policies from expert rules, jointly optimizing the target policy and reuse probabilities under a constrained Markov decision process (CMDP). This fusion of heterogeneous priors accelerates convergence and enhances robustness. Constrained stochastic successive convex approximation (CSSCA) handles non-convex objectives and constraints, while a critic trained from mixed offline-online data improves sample efficiency by reusing pre-collected experience. We prove almost-sure convergence to a Karush-Kuhn-Tucker (KKT) point. Simulations on delay-constrained multi-user multiple-input multiple-output (MU-MIMO) power control and Cramer-Rao bound (CRB)-constrained multiple-input multiple-output integrated sensing and communication (MIMO-ISAC) beamforming demonstrate that Fused-CPRO improves empirical performance and converges substantially faster than representative baselines.

View free PDFSource page

Related papers

arxiveess.SP2026-07-22

WARA: A Closed-Loop Multi-Agent Framework for Wireless Optimization Autoresearch

Yuan Guo, Yilong Chen, Chao Hu, Xianghao Yu, Liang Hong, Jie Xu

Large language model (LLM) agents have shown growing capabilities in tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific research. To the best of our knowledge, this paper presents the first end-to-end autore…

View free PDFSource page
arxiveess.SP2026-07-09

Beyond Single-Band: Analysis and Resource Allocation for Multi-band ISAC Systems

Haotian Liu, Zhiqing Wei, Xingwang Li, Yunxin Geng, Qixun Zhang, Zhiyong Feng

Integrated sensing and communication (ISAC) has emerged as a pivotal technology for sixth-generation wireless networks to empower high-precision sensing. The demand for superior sensing resolution and the reality of spectrum fragmentation have driven the research of multi-band IS…

View free PDFSource page
arxiveess.SP2026-07-14

Positional Attention-based Graph Neural Network for Learning Permutation Non-equivariant Wireless Policies

Baichuan Zhao, Chenyang Yang, Jianyu Zhao, Di Zhang

Graph neural networks (GNNs) have emerged as a promising approach to learning wireless policies efficiently by leveraging topology prior and incorporating relational inductive biases. However, when the optimal policy is not permutation equivariant (PE), conventional GNNs suffer f…

View free PDFSource page
arxiveess.SP2026-07-17

Energy-Efficient Resource Allocation for Six-Dimensional Movable Antenna Systems

Ziyun Zhang, Ruotong Zhao, Shaokang Hu, Derrick Wing Kwan Ng

This paper investigates the energy-efficiency (EE) maximization problem for a multiuser wireless network equipped with six-dimensional movable antennas (6DMAs), where the three-dimensional (3D) positions and orientations of the antennas are jointly optimized to fully exploit the…

View free PDFSource page