overfeed.news

Faulty reward functions in the wild

OpenAI News

9y

Age

Published
Collected

Reinforcement learning algorithms can break in surprising, counterintuitive ways. In this post we’ll explore one failure mode, which is where you misspecify your reward function.

Read the full article at openai.com

overfeed.news indexes and links. We publish a short excerpt — the full article stays at OpenAI News.

More from OpenAI News

Log in to follow this source
Faulty reward functions in the wild — overfeed.news