Risks from Learned Optimization in Advanced Machine Learning Systems

E2162383 UNEXPLORED

"Risks from Learned Optimization in Advanced Machine Learning Systems" is an influential AI safety paper that analyzes how powerful machine learning models can develop internal optimization processes (mesa-optimizers) whose objectives may diverge from those intended by their designers, posing novel alignment risks.

Try in SPARQL Jump to: Surface forms Referenced by

All labels observed (1)

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

Evan Hubinger coAuthorOf Risks from Learned Optimization in Advanced Machine Learning Systems