overfeed.news

Seção

Notícias

Manchetes das redações e blogs que acompanhamos.

2.476documentos

OpenAI News

Detecting misbehavior in frontier reasoning models

Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.

en
Anúncio

Assine o Pro, sem anúncios

Faça upgrade para uma leitura sem interrupções e acesso prioritário a novas fontes.

Ver preços
OpenAI News

Deep research System Card

This report outlines the safety work carried out prior to releasing deep research including external red teaming, frontier risk evaluations according to our Preparedness Framework, and an overview of the mitigations we built in to address key risk areas.

en