Reinforcement learning with prediction-based rewards
October 31, 20182 views1 min read
We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.
Share:
Источник: OpenAI Blog
Related News
AI
Neural NetworksOpenAI Releases GPT-5
OpenAI released GPT-5, a new AI model with improved capabilities in text generation, mathematics, and programming.
5h ago66
AI
Neural NetworksCursor capitalizes on GitHub frustration, launches rival hosting platform
6h ago19
Neural NetworksRobin Williams’ Instagram account brought back to fight ‘AI abuse’
9h ago16