Triple
T18629602
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | Arthur Guez |
E455377
|
entity |
| Predicate | coAuthorOf |
P2389
|
FINISHED |
| Object | Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search |
—
|
NE NERFINISHED |
How this triple was built (3 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search | Statement: [Arthur Guez, coAuthorOf, Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search]
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search Context triple: [Arthur Guez, coAuthorOf, Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search]
-
A.
Practical Bayesian Optimization of Machine Learning Algorithms
Practical Bayesian Optimization of Machine Learning Algorithms is a seminal research paper that introduced efficient Bayesian optimization techniques for automatically tuning hyperparameters of complex machine learning models.
-
B.
V-trace off-policy correction algorithm
The V-trace off-policy correction algorithm is a method for stabilizing and improving learning in distributed deep reinforcement learning by correcting for discrepancies between behavior and target policies.
-
C.
Monte Carlo tree search
Monte Carlo tree search is a heuristic search algorithm that uses random sampling of game states to build and explore a search tree, enabling strong decision-making in complex domains like Go and other board games.
-
D.
Deterministic policy gradient algorithms
Deterministic policy gradient algorithms are a class of reinforcement learning methods that learn policies with deterministic actions in continuous action spaces by directly optimizing expected returns via gradient-based updates.
-
E.
Natural Policy Gradient
Natural Policy Gradient is a reinforcement learning optimization method that improves policy gradient updates by accounting for the geometry of the parameter space using the Fisher information matrix, leading to more stable and efficient learning.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search Target entity description: "Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search" is a research paper that introduces a scalable, sample-based planning method for Bayes-adaptive reinforcement learning, enabling more efficient decision-making under model uncertainty.
-
A.
Practical Bayesian Optimization of Machine Learning Algorithms
Practical Bayesian Optimization of Machine Learning Algorithms is a seminal research paper that introduced efficient Bayesian optimization techniques for automatically tuning hyperparameters of complex machine learning models.
-
B.
V-trace off-policy correction algorithm
The V-trace off-policy correction algorithm is a method for stabilizing and improving learning in distributed deep reinforcement learning by correcting for discrepancies between behavior and target policies.
-
C.
Monte Carlo tree search
Monte Carlo tree search is a heuristic search algorithm that uses random sampling of game states to build and explore a search tree, enabling strong decision-making in complex domains like Go and other board games.
-
D.
Deterministic policy gradient algorithms
Deterministic policy gradient algorithms are a class of reinforcement learning methods that learn policies with deterministic actions in continuous action spaces by directly optimizing expected returns via gradient-based updates.
-
E.
Natural Policy Gradient
Natural Policy Gradient is a reinforcement learning optimization method that improves policy gradient updates by accounting for the geometry of the parameter space using the Fisher information matrix, leading to more stable and efficient learning.
- F. None of above. chosen
Provenance (2 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69d8d38cc7948190a55ea64e5638994e |
completed | April 10, 2026, 10:40 a.m. |
| NER | Named-entity recognition | batch_69e54f06f4a081909b64f33814577488 |
completed | April 19, 2026, 9:54 p.m. |
Created at: April 10, 2026, 11:46 a.m.