V-trace off-policy correction algorithm

E1276524 UNEXPLORED

The V-trace off-policy correction algorithm is a method for stabilizing and improving learning in distributed deep reinforcement learning by correcting for discrepancies between behavior and target policies.

All labels observed (1)

Label Occurrences
V-trace off-policy correction algorithm canonical 1

How this entity was disambiguated

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

IMPALA notableComponent V-trace off-policy correction algorithm