Reinforcement Learning Primer
Multi-Agent Reinforcement Learning Primer — Optimizing in a World Where the Other Side Is Learning Too
When multiple agents learn at the same time, a single-agent MDP no longer holds. This article organizes non-stationarity, cooperative/competitive/mixed settings, CTDE, the credit-assignment problem, and MADDPG and QMIX, with equations and diagrams.