overfeed.news

Detecting and reducing scheming in AI models

OpenAI Newsen

11mês

Idade

Publicado
Coletado

Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. The team shared concrete examples and stress tests of an early method to reduce scheming.

Leia o artigo completo em openai.com

O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em OpenAI News.

Mais de OpenAI News

Entre para seguir esta fonte
Detecting and reducing scheming in AI models — overfeed.news