Improving instruction hierarchy in frontier LLMs
March 10, 20262 views1 min read
IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety steerability, and resistance to prompt injection attacks.
Share:
Источник: OpenAI Blog
Related News
AI
Neural NetworksOpenAI Releases GPT-5
OpenAI released GPT-5, a new AI model with improved capabilities in text generation, mathematics, and programming.
4h ago66
AI
Neural NetworksCursor capitalizes on GitHub frustration, launches rival hosting platform
6h ago17
Neural NetworksRobin Williams’ Instagram account brought back to fight ‘AI abuse’
8h ago12