overfeed.news

Improving instruction hierarchy in frontier LLMs

OpenAI News

5mo

Age

Published
Collected

IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety steerability, and resistance to prompt injection attacks.

Read the full article at openai.com

overfeed.news indexes and links. We publish a short excerpt — the full article stays at OpenAI News.

More from OpenAI News

Log in to follow this source
Improving instruction hierarchy in frontier LLMs — overfeed.news