Ankündigung des Herstellerserreichbar
Video generation models as world simulators
Der technische Bericht vom 15.02.2024, in dem OpenAI ein Videomodell ausdrücklich als Weg zu einem Simulator der physischen Welt beschreibt. Er nennt die Bauform (Diffusion über Raum-Zeit-Stücke statt über Text-Token) und zählt die Fähigkeiten auf, die dabei ohne eigenen Bauteil entstanden sind. Der Bericht sagt selbst, dass Modell- und Umsetzungsdetails fehlen, und er nennt die Grenzen beim Namen.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 05.09.2026:
Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.
bestätigt 24.09.2026Whereas LLMs have text tokens, Sora has visual patches.
bestätigt 24.09.2026We find that video models exhibit a number of interesting emergent capabilities when trained at scale. These capabilities enable Sora to simulate some aspects of people, animals and environments from the physical world.
bestätigt 24.09.2026Sora currently exhibits numerous limitations as a simulator. For example, it does not accurately model the physics of many basic interactions, like glass shattering.
bestätigt 24.09.2026Model and implementation details are not included in this report.
bestätigt 24.09.2026
Text-zu-Video im GlossarZeitliche Beständigkeit im Glossar09 Weltmodelle13 Video, Musik und 3D