Date of Award

6-26-2026

Date Published

August 2026

Degree Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

Department

Electrical Engineering and Computer Science

Advisor(s)

Qinru Qiu

Keywords

artificial intelligence;belief representation;communication learning;federated reinforcement learning;multi-agent reinforcement learning;reinforcement learning

Abstract

Reinforcement learning is a powerful approach for training cooperative multi-agent systems. However, it is usually studied under idealized conditions. Standard formulations assume abundant experience, reliable communication, and cleanly observable state. Physical multi-agent systems rarely satisfy these assumptions, and instead operate under several kinds of scarcity. This dissertation studies reinforcement learning for cooperative multi-agent systems under two practical resource constraints. We show that effective learning in these settings depends not only on how much communication, computation, or experience is available, but also on how limited information and experience are represented, reused, and evaluated. Part I addresses information constraints in cooperative multi-agent reinforcement learning (MARL), where each agent must coordinate from a partial view of a global state that it cannot directly observe. Belief-map Assisted Multi-agent System (BAMS) trains each agent to decode a structured belief over the global state from its local observations and received messages. The auxiliary belief decoder is supervised by privileged information during training and removed at execution, while the shaped hidden representation remains part of the decentralized policy. Belief-based Predictive Auxiliary Learning (BEPAL) extends this idea with a flexible predictive decoder whose targets include task-relevant present and future dynamics, and it improves coordination across cooperative navigation, traffic, team-play, and warehouse-logistics environments. We then propose a communication unlearning pipeline that suppresses redundant messages from a trained policy while preserving task performance, which reduces the communication overhead that can emerge from belief-supervised communication policies. Part II addresses experience constraints in federated reinforcement learning for UAV teams in hazardous environments, where the rate of safe and informative experience is bounded by mission safety rather than by computation. We propose the Experience-Constrained Hierarchical Federated RL (EC-HFRL) framework and three diagnostics: the Union Coverage Ratio (UCR), the Key Enrichment Ratio (KER), and the Key TD Contribution (KTC). Under a fixed experience budget, increasing learner participation mainly changes how a bounded experience pool is reused rather than expanding it. Moreover, performance is strongly associated with whether task-relevant key experiences are admitted into the high-error replay subset. Higher participation mainly raises energy cost. We further introduce an anisotropic reward-shaping design and analyze how key-experience composition shifts across training stages in two hazard regimes. Across both parts, the same lesson appears from different directions: under practical constraints, performance depends strongly on how limited information and experience are structured and used, not only on how much communication, computation, or learner participation is available. Part I uses auxiliary belief supervision to improve the representations agents form from partial observations and communication, and then reduces communication that carries little marginal value. Part II shows that, under a fixed experience budget, learning performance is closely associated with which experiences enter and dominate the replay-based learning signal, and it introduces diagnostics that make this composition measurable.

Access

Open Access

Share

COinS