All News
Neural Networks

Detecting misbehavior in frontier reasoning models

March 10, 20252 views1 min read

Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.

Want to try AI?

Compare the best neural networks in one place — for free

Go to neural networks
Share:
Источник: OpenAI Blog

Related News