policy gradient theorem
E1874121
UNEXPLORED
The policy gradient theorem is a fundamental result in reinforcement learning that provides a way to compute the gradient of expected return with respect to policy parameters, enabling gradient-based optimization of stochastic policies.
All labels observed (1)
| Label | Occurrences |
|---|---|
| policy gradient theorem canonical | 1 |
Referenced by (1)
Full triples — surface form annotated when it differs from this entity's canonical label.