Ankündigung des Herstellerserreichbar
Gemini Robotics brings AI into the physical world
deepmind.google (externe Seite)
Google DeepMind stellt am 12.03.2025 zwei Modelle vor und führt dabei die Bauform Vision-Language-Action als eigene Kategorie ein: ein Sprachmodell, das Bewegung als zusätzliche Ausgabeform bekommt. Nennt Apptronik als Partner für humanoide Roboter, den Datensatz ASIMOV für Sicherheitsmessungen und eine Verdopplung auf einem Verallgemeinerungs-Maßstab. Diese Zahl stammt aus dem eigenen technischen Bericht.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 07.09.2026:
an advanced vision-language-action (VLA) model that was built on Gemini 2.0 with the addition of physical actions as a new output modality for the purpose of directly controlling robots
bestätigt 24.09.2026So far however, those abilities have been largely confined to the digital realm.
bestätigt 24.09.2026on average, Gemini Robotics more than doubles performance on a comprehensive generalization benchmark compared to other state-of-the-art vision-language-action models
bestätigt 24.09.2026The physical safety of robots and the people around them is a longstanding, foundational concern in the science of robotics.
bestätigt 24.09.2026