All News
Neural Networks

Detecting and reducing scheming in AI models

September 17, 20250 views1 min read

Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. The team shared concrete examples and stress tests of an early method to reduce scheming.

Want to try AI?

Compare the best neural networks in one place — for free

Go to neural networks
Share:
Источник: OpenAI Blog

Related News