Improving mathematical reasoning with process supervision
3a
Idade
- Publicado
- Coletado
We’ve trained a model to achieve a new state-of-the-art in mathematical problem solving by rewarding each correct step of reasoning (“process supervision”) instead of simply rewarding the correct final answer (“outcome supervision”). In addition to boosting performance relative to outcome supervision, process supervision also has an important alignment benefit: it directly trains the model to produce a chain-of-thought that is…
O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em OpenAI News.
Mais de OpenAI News
Entre para seguir esta fonteBetter answers, broader thinking: What students gain from ChatGPT and critical-thinking training
A randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.
Expanding OpenAI’s presence in Brazil
OpenAI is expanding its presence in Brazil, deepening engagement with developers, businesses, and communities to support AI adoption across the country.
Learning never stops: How AI makes learning continuous
OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
Bringing ChatGPT for Teachers to more U.S. school districts
ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.