Continuous control with deep reinforcement learning
E1284937
UNEXPLORED
"Continuous control with deep reinforcement learning" is a highly influential research paper that introduced deep neural network methods for solving continuous-action reinforcement learning tasks, notably using deterministic policy gradients.
All labels observed (1)
| Label | Occurrences |
|---|---|
| Continuous control with deep reinforcement learning canonical | 3 |
How this entity was disambiguated
This entity first appeared as the object of triple T17738637 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Continuous control with deep reinforcement learning Context triple: [Timothy P. Lillicrap, coAuthorOf, Continuous control with deep reinforcement learning]
-
A.
Deep Q-Learning
Deep Q-Learning is a reinforcement learning algorithm that uses deep neural networks to approximate Q-values, enabling agents to learn effective policies directly from high-dimensional inputs like raw images.
-
B.
Natural Policy Gradient
Natural Policy Gradient is a reinforcement learning optimization method that improves policy gradient updates by accounting for the geometry of the parameter space using the Fisher information matrix, leading to more stable and efficient learning.
-
C.
Atari deep Q-network
The Atari deep Q-network is a pioneering deep reinforcement learning system that learned to play a wide range of Atari 2600 video games directly from raw pixels at human-level or better performance.
-
D.
Soft Actor-Critic
Soft Actor-Critic is a model-free deep reinforcement learning algorithm that combines off-policy learning with entropy maximization to achieve stable and sample-efficient continuous control.
-
E.
Proximal Policy Optimization
Proximal Policy Optimization is a popular reinforcement learning algorithm that improves policy gradient methods by using clipped objective functions to achieve stable and efficient training.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Continuous control with deep reinforcement learning Target entity description: "Continuous control with deep reinforcement learning" is a highly influential research paper that introduced deep neural network methods for solving continuous-action reinforcement learning tasks, notably using deterministic policy gradients.
-
A.
Deep Q-Learning
Deep Q-Learning is a reinforcement learning algorithm that uses deep neural networks to approximate Q-values, enabling agents to learn effective policies directly from high-dimensional inputs like raw images.
-
B.
Natural Policy Gradient
Natural Policy Gradient is a reinforcement learning optimization method that improves policy gradient updates by accounting for the geometry of the parameter space using the Fisher information matrix, leading to more stable and efficient learning.
-
C.
Atari deep Q-network
The Atari deep Q-network is a pioneering deep reinforcement learning system that learned to play a wide range of Atari 2600 video games directly from raw pixels at human-level or better performance.
-
D.
Soft Actor-Critic
Soft Actor-Critic is a model-free deep reinforcement learning algorithm that combines off-policy learning with entropy maximization to achieve stable and sample-efficient continuous control.
-
E.
Proximal Policy Optimization
Proximal Policy Optimization is a popular reinforcement learning algorithm that improves policy gradient methods by using clipped objective functions to achieve stable and efficient training.
- F. None of above. chosen
Referenced by (3)
Full triples — surface form annotated when it differs from this entity's canonical label.