The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
April 19, 20242 views1 min read
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.
Share:
Источник: OpenAI Blog
Related News
AI
Neural NetworksOpenAI Releases GPT-5
OpenAI released GPT-5, a new AI model with improved capabilities in text generation, mathematics, and programming.
5h ago68
AI
Neural NetworksCursor capitalizes on GitHub frustration, launches rival hosting platform
7h ago19
Neural NetworksRobin Williams’ Instagram account brought back to fight ‘AI abuse’
9h ago16