overfeed.news

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

OpenAI News

2y

Age

Published
Collected

Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.

Read the full article at openai.com

overfeed.news indexes and links. We publish a short excerpt — the full article stays at OpenAI News.

More from OpenAI News

Log in to follow this source
The Instruction Hierarchy: Training LLMs to Prioritize…