Video generation models as world simulators
2y
- Published
- Collected
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes. Our largest model, Sora, is capable of generating a minute of high fidelity video. Our results…
overfeed.news indexes and links. We publish a short excerpt — the full article stays at OpenAI News.
More from OpenAI News
Log in to follow this sourceAsana cuts model costs 76x in browser tests with GPT-6.1 Sol
Using GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
Sophos cuts threat investigation time by 96% with OpenAI Daybreak
Discover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight.
How Oracle turns days of work into minutes with ChatGPT and Codex
Across recruiting, engineering, and operations, Oracle turns specialist knowledge into fast, repeatable workflows with ChatGPT Work and Codex.
LegalOn halves Codex costs while maintaining development speed
LegalOn cut estimated daily Codex costs by 65% while maintaining development speed. It matched Astra, Sol, and Luna to tasks and managed budgets strategically.