Proximal Policy Optimization
July 20, 20172 views1 min read
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance.
Share:
Источник: OpenAI Blog
Related News
AI
Neural NetworksOpenAI Releases GPT-5
OpenAI released GPT-5, a new AI model with improved capabilities in text generation, mathematics, and programming.
4h ago66
AI
Neural NetworksCursor capitalizes on GitHub frustration, launches rival hosting platform
6h ago15
Neural NetworksRobin Williams’ Instagram account brought back to fight ‘AI abuse’
8h ago12