Odyssey-3 is a foundation world model that generates embodied environments from a prompt and predicts in real time how those environments change as a person or an agent takes actions or introduces events, using previous observations and the latest inputs. It is described as a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time. The model learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations, and developers use that knowledge both to simulate environments and to train policies for different physical systems. The research preview is available now, and physical AI developers who want to build with Odyssey-3 are invited to get in touch for access.
The team behind Odyssey-3 comes from a decade spent building driverless cars, where predicting the world was essential to determining what a car should do next. The founders started Odyssey to pursue that idea far beyond the roads and to build a general-purpose technology that could bring learned world knowledge to all machines and tasks. Physical accuracy has been made a central focus of the research, on the reasoning that a foundation model for physical intelligence must learn to predict how the world actually behaves. In the team's own framing, with Odyssey-3 that idea is now being realized.
Odyssey-3 generates embodied environments in real time and predicts how they change as a person or agent takes actions or introduces events. You can move through the environment or introduce an event during generation and observe how the model responds. The current preview provides first-person and third-person navigation alongside independent camera movement, giving different ways to interact with and inspect the model's predictions. The stated design goal was to enable dynamic, open-ended interactions with an environment that responds as the user acts.
Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified's video-to-video benchmark, achieving 66.1, the highest reported score. Physics-IQ, a benchmark from Anates Labs and DeepMind, tests physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics by asking models to continue videos of real physical experiments and comparing their predictions with what actually happened. Odyssey-3 Pro also scores 54.7 in image-to-video. Reported video-to-video scores are 51.8 with base prompts, 61.6 with prompt enhancement, and 64.4 best-of-8 for the Odyssey-3 480p series, and 63.4 with prompt enhancement and 66.1 best-of-8 for Odyssey-3 Pro 720p. On image-to-video, Odyssey-3 records 41.0 with base prompts, 48.8 with prompt enhancement, and 52.8 best-of-8, while Odyssey-3 Pro records 50.0 with prompt enhancement and 54.7 best-of-8. Odyssey-3 also improves the measured tradeoff between physical accuracy and generation cost, making it possible to generate more simulations within the same compute budget. Resolutions are 832x480 for Odyssey-3 and 1280x720 for Pro.
WorldMark measures control-following, visual quality, and world memory. In Odyssey's evaluation, using the benchmark's own captions and the mean of its 13 reported metric scores, Odyssey-3 ranks first in first-person stylized environments with 77.2, third-person real environments with 79.0, and third-person stylized environments with 76.3, and places third in first-person real environments with 80.6. The company notes that these results measure specific properties of generated worlds, and that applying the model to a physical system also requires evaluating the behaviors that matter for that machine and its tasks.
Odyssey-3's learned world knowledge can be applied to different systems by training an action decoder or policy on paired observations and actions. These learned components translate that knowledge into the controls required by a particular machine, letting developers adapt the foundation model to a new body or task. With only tens of hours of robot demonstrations, Odyssey-3 completed manipulation tasks such as "Pour the cereal into the bowl" and "Close the screwbox" and showed recovery behaviors absent from those demonstrations, including reorienting a gripper after a missed grasp and retrieving a dropped object in an unusual position. Flexion has built humanoid control policies on Odyssey-3; the resulting policies exceeded the performance of the tested VLA baselines under environmental changes and continued to perform tasks under lighting changes that caused those baselines to fail.
Odyssey-3 was also adapted to drive a car on real roads in India, training a driving policy on just 20 hours of driving data while keeping the Odyssey-3 backbone frozen. The policy uses the model's visual representations to predict waypoints ahead of the car, allowing it to drive in closed loop, and was demonstrated with instructions such as "Take the first roundabout exit" and "Drive along the road". Separately, Odyssey-3 was adapted to generate observations for particular sensor arrangements: in an early experiment using the front three cameras of an autonomous-driving dataset, an Odyssey-3 training checkpoint produced driving sequences with three camera views generated together after just 100 training steps.
Odyssey-3 can also help train agents. An agent is an AI system that pursues a goal by observing its surroundings, choosing actions, and using what happens to decide what to do next. A world model can provide the environment in which those decisions are made, giving a way to study how an agent responds to changing conditions and whether it can complete a task inside a world whose behavior is learned. In Odyssey's task-completion demonstration, an agent receives a natural-language goal and pursues it inside Odyssey-3, observing the generated world as it works toward the task.
Odyssey-3's training data combines internet video with time-localized, schema-verified event annotations, gameplay recordings with time-aligned keyboard and mouse inputs, and simulated rigid-body interactions with captions and metadata. Together these sources connect diverse observations with descriptions of what happens and, where available, the actions that produced it. The model is built as a multi-step video diffusion transformer, using temporally resolved prompts and controls to guide how the world unfolds. It is then extended autoregressively through teacher forcing and causal masking, training it to continue from preceding observations and predict future states conditioned on action inputs. Finally, a post-training pipeline combines distribution-matching and adversarial distillation to produce a distilled variant of Odyssey-3, a few-step model capable of real-time interaction.
Odyssey believes world models will power increasingly capable physical AI, generate environments in which other intelligences can train, and enable new kinds of human experiences. Developers building robots, humanoids, self-driving cars, drones, or any other autonomous system are invited to explore how foundation world models can accelerate their work. The research preview lets people prompt an environment, act within it, and see how the world model responds, while API access is available by getting in touch with the team.
In summary, Odyssey-3 is presented as the company's most powerful foundation world model, combining real-time, prompt-driven environment generation with strong physical accuracy results on Physics-IQ Verified and WorldMark, and with demonstrated adaptation to robot arms, humanoids, vehicles, multi-sensor data generation, and agent training.