Learning Montezuma’s Revenge from a single demonstration
8a
Idade
- Publicado
- Coletado
We’ve trained an agent to achieve a high score of 74,500 on Montezuma’s Revenge from a single human demonstration, better than any previously published result. Our algorithm is simple: the agent plays a sequence of games starting from carefully chosen states from the demonstration, and learns from them by optimizing the game score using PPO, the same reinforcement learning algorithm…
O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em OpenAI News.
Mais de OpenAI News
Entre para seguir esta fonteBetter answers, broader thinking: What students gain from ChatGPT and critical-thinking training
A randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.
Expanding OpenAI’s presence in Brazil
OpenAI is expanding its presence in Brazil, deepening engagement with developers, businesses, and communities to support AI adoption across the country.
Learning never stops: How AI makes learning continuous
OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
Bringing ChatGPT for Teachers to more U.S. school districts
ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.