Reinforcement Learning Primer
Policy Gradients, PPO, and SAC — Stable Continuous Control for Robots
Policy gradients update a policy directly so that steering, thrust, and joint torque can remain continuous. This article connects Actor-Critic, PPO, and SAC, then covers clipping, entropy regularization, evaluation, and the safety boundary for Sim-to-Real transfer.