overfeed.news

Faulty reward functions in the wild

OpenAI Newsen

9a

Idade

Publicado
Coletado

Reinforcement learning algorithms can break in surprising, counterintuitive ways. In this post we’ll explore one failure mode, which is where you misspecify your reward function.

Leia o artigo completo em openai.com

O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em OpenAI News.

Mais de OpenAI News

Entre para seguir esta fonte
Faulty reward functions in the wild — overfeed.news