explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What Odyssey-3 is
  • The Physics-IQ Verified leaderboard
  • The cost math is as interesting as the score
  • WorldMark: first in three of four splits, by Odyssey's own count
  • From video model to robot, car and agent
  • How it was built
  • What is confirmed and what is not
  • What this means for what you build
  • Related reading on explainx.ai
← Back to blog

explainx / blog

Odyssey-3 Tops Physics-IQ Verified at 66.1: World Model Results

World Models, Physical AI, Robotics, Benchmarks, AI News

Odyssey-3 Pro scores 66.1 on Physics-IQ Verified and ranks first in 3 of 4 WorldMark splits. The leaderboard, cost math, robot demos and what to verify.

Oct 9, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Odyssey-3 Tops Physics-IQ Verified at 66.1: World Model Results

On October 8, 2026, Odyssey launched Odyssey-3, which it calls its most powerful foundation world model. The headline claim is a new state of the art on Physics-IQ Verified, a benchmark from Anates Labs and Google DeepMind: Odyssey-3 Pro reaches 66.1 on the video-to-video track, the highest reported score. Odyssey also reports first place in three of four WorldMark categories in its own evaluation, and shows the model adapted to robot arms, humanoids and a car driving on real roads.

We covered the earlier Odyssey-3 introduction in September, which had no numbers. This post covers what the October launch adds: the leaderboard, the cost-versus-accuracy math, the adaptation demos, and what to check before you rely on the claims. The source is Odyssey's launch post and the leaderboard graphic it published, dated October 7, 2026.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?A foundation world model: an autoregressive diffusion transformer that predicts how scenes evolve
Headline score?66.1 on Physics-IQ Verified video-to-video (Odyssey-3 Pro, best-of-8)
Without best-of-8?63.4 for Pro; 61.6 for base Odyssey-3 with prompt enhancement
Image-to-video?54.7 for Pro with best-of-8
Main rival on the board?FLUX 3 large from Black Forest Labs at 64.35 with best-of-8
Where is the catch?Best-of-N sampling, cost figures that exclude prompt-rewriting fees, and company-run WorldMark tests
Can I try it?A free research preview, with API access by contacting Odyssey
Is it a robot brain?It is a foundation for one; the robot demos used small adaptation datasets

What Odyssey-3 is

Odyssey describes Odyssey-3 as "a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time." In plain terms: it is a video model that continues a scene frame by frame, conditioned on what happened before and on actions or events you introduce. Developers use that learned knowledge two ways, to simulate environments and to train policies for physical systems.

The research preview shows first-person and third-person navigation plus independent camera movement. You prompt an environment, move through it, or introduce an event, and the model predicts how the world responds in real time. The real-time variant is a distilled few-step model, discussed in the training section below.

For background on the category, see our explainer on what world models are and the Odyssey Agora-2 multi-agent world model.

The Physics-IQ Verified leaderboard

Physics-IQ asks models to continue videos of real physical experiments, covering fluid dynamics, optics, solid mechanics, magnetism and thermodynamics, and compares the prediction with what actually happened. The leaderboard graphic Odyssey shared, for the video-to-video, verified-score, all-sampling, all-models view, ranks 11 entries:

table · 6 cols
RankModelScoreSamplingOutput FPSCost per video
1Odyssey-3 Pro66.10%Best of 816$3.484
2Odyssey-364.43%Best of 816$1.430
3FLUX 3 large (Black Forest Labs)64.35% ± 0.22Best of 824$16.430
4Odyssey-3 Pro63.37% ± 0.63Single16$0.444
5Odyssey-361.56% ± 1.20Single16$0.187
6FLUX 3 large61.11% ± 0.55Single24$2.080
7Physical Registry + PRS (Microsoft Research Asia)52.61%Best of 424not disclosed
8Odyssey-3, no LLM, base prompts51.76% ± 0.69Single16$0.177
9Cosmos3 Super (NVIDIA, open source)50.80% ± 2.20Single24$0.823
10Physical Registry (Microsoft Research Asia)50.17% ± 0.72Single24not disclosed
11Cosmos3 Nano (NVIDIA, open source)43.00% ± 2.00Single24$0.823

Microsoft Research Asia's entries are marked "Not Yet Available." The "±" is sample standard deviation; best-of-8 rows have none because they use one run. All Odyssey and FLUX rows use LLM prompt enhancement with custom prompts, except the Odyssey-3 row at 51.76, which uses base prompts and no LLM.

Odyssey also reports 54.7 on image-to-video for Odyssey-3 Pro with best-of-8, with Odyssey-3 at 52.8 best-of-8, 48.8 with prompt enhancement and 41.0 on base prompts. Its image-to-video chart compares against a long list of video generators including Seedance 2.5, MiniMax H3, Gemini Omni Flash, Veo 3.1, Grok Imagine Video, Sora 2, Wan 2.2 and others, but the numbers for those were not legible in the text we have, so we do not repeat them.

The cost math is as interesting as the score

Odyssey argues that Odyssey-3 "improves the measured tradeoff between physical accuracy and generation cost," meaning more simulations for the same compute budget. Using the table above:

table · 3 cols
ComparisonScore gapCost ratio
Odyssey-3, best-of-8 (64.43) vs FLUX 3 large, best-of-8 (64.35)about equalFLUX costs about 11.5x more ($16.43 vs $1.43)
Odyssey-3 Pro, single (63.37) vs FLUX 3 large, single (61.11)+2.26 pointsFLUX costs about 4.7x more ($2.08 vs $0.444)
Odyssey-3 Pro best-of-8 (66.10) vs Odyssey-3 best-of-8 (64.43)+1.67 pointsPro costs about 2.4x more
Odyssey-3 with prompt enhancement (61.56) vs base prompts (51.76)+9.8 points$0.187 vs $0.177

Two points deserve scrutiny.

Prompt enhancement does most of the work, and its fee is not in the cost. Going from base prompts to LLM-enhanced prompts adds about ten points for about a cent, but Odyssey's footnote says its cost "assumes $1 per MI355X GPU-hour, excluding prompt-rewriting fees." The comparison models' costs "include prompt fees where reported." That is an apples-to-oranges cost comparison, and the real Odyssey cost is higher than shown, by an amount the footnote does not give.

Best-of-N is a different kind of result. Best-of-8 generates eight samples and keeps the best one. The footnote says scores average four runs, while "best-of-8 uses 1 run with the same prompts." Best-of-N needs a way to pick the winner, and in a real application you may not have ground truth to select against. Rows without best-of-N are the fairer guide to what you get from one generation: 63.4 for Pro and 61.6 for base Odyssey-3, against 61.1 for FLUX 3 large.

Other footnotes: Odyssey-3 renders at 832×480 and Pro at 1280×720, Odyssey runs at 16 frames per second against 24 for FLUX 3 and Cosmos3, and Cosmos3's video-to-video cost uses compute-based pricing from October 1. All of that affects what "cost per video" means.

WorldMark: first in three of four splits, by Odyssey's own count

WorldMark measures control-following, visual quality and world memory. Odyssey evaluated models using the benchmark's own captions and the mean of its 13 reported metrics, so these are Odyssey-run numbers.

table · 3 cols
SplitOdyssey-3Closest competitor
First-person stylized77.2 (1st)LingBot-World 77.0, AlayaWorld 76.7
First-person real80.6 (3rd)Lyra 2.0 84.4, AlayaWorld 83.0
Third-person real79.0 (1st)HY-World 1.5 76.9
Third-person stylized76.3 (1st)HY-World 1.5 75.1

The first-person stylized lead is 0.2 points, which is within the range where evaluation noise could matter. Odyssey itself cautions that these results measure specific properties of generated worlds, and that applying the model to a physical system requires evaluating the behaviors that matter for that machine.

From video model to robot, car and agent

The most strategic part of the launch is the claim that one world model can be adapted to different bodies by training a small action decoder or policy on paired observations and actions.

table · 3 cols
DemoWhat Odyssey saysWhat to keep in mind
Robot armsWith only tens of hours of demonstrations, completed tasks such as pouring cereal and closing a screwbox, and showed recovery behaviors absent from the demonstrationsSmall set of tasks; no success rates published in the post
HumanoidsFlexion built humanoid policies on Odyssey-3 that exceeded tested VLA baselines under environmental changes and kept working under lighting changes that broke those baselinesThird-party claim relayed by Odyssey; baselines unspecified
DrivingA policy trained on 20 hours of driving data in India, with the backbone frozen, predicts waypoints and drives in closed loopEarly result; no safety or intervention metrics given
Multi-sensorA training checkpoint produced three-camera driving sequences after 100 training stepsAn early experiment
Agent trainingAn agent pursues a natural-language goal inside the generated worldA demonstration, not a benchmark

These are credible directions and a useful pitch, since data for physical systems is scarce. They are not independent evidence that the policies are robust. For comparison with another physical-AI stack, see NVIDIA Cosmos 3 and NVIDIA's SIGGRAPH Cosmos update, and for robot hardware with a physical API, 1X NEO.

How it was built

Odyssey's training data combines three sources: internet video with time-localized, schema-verified event annotations, gameplay recordings with time-aligned keyboard and mouse inputs, and simulated rigid-body interactions with captions and metadata. The goal is to connect observations with descriptions of what happens and, where available, the actions that caused it.

The recipe has three stages:

  1. A multi-step video diffusion transformer that uses temporally resolved prompts and controls to guide how the world unfolds.
  2. Autoregressive extension through teacher forcing and causal masking, so the model continues from earlier observations and predicts future states given actions.
  3. Post-training with distribution-matching and adversarial distillation, producing a few-step distilled variant fast enough for real-time interaction.

That last step is why "Flash" can run live while "Pro" is the quality tier.

What is confirmed and what is not

table · 2 cols
ClaimStatus
Physics-IQ Verified scores and rankingFrom a leaderboard graphic dated October 7, 2026; benchmark from Anates Labs and Google DeepMind; submission and verification process not detailed in our sources
CostsOdyssey's assumptions; exclude its prompt-rewriting fees
WorldMark ranksOdyssey-run, using the benchmark's captions and mean of 13 metrics
Robot, humanoid, driving demosOdyssey and partner claims, with small data budgets
AvailabilityResearch preview; API by contact

What this means for what you build

  1. If you build simulators or synthetic-data pipelines, try the preview and measure physical plausibility on your own scenes, not only the benchmark. The cost per video matters when you need thousands of rollouts.
  2. If you train robot policies, the interesting part is the adapter recipe: a frozen backbone plus a small decoder trained on tens of hours. Reproduce it on your hardware before assuming the headline holds.
  3. Read the sampling column. Compare single-generation scores with single-generation scores. Best-of-N is a tool for offline data generation, not a free accuracy boost.
  4. Ask for the full cost. Include prompt rewriting, resolution and frame rate when you compare per-video prices.
  5. Watch for independent replication. A benchmark leaderboard run by outside organizations is stronger evidence than a vendor's chart, and the other world models in this space, including Runway's GWM Worlds 2, ByteDance's Seedance world model and Tencent's HY-World 2, will be measured against the same bars.

Related reading on explainx.ai

  • Odyssey-3 introduction: a unified foundation world model
  • Odyssey Agora-2: a multi-agent world model for 20 players
  • What are world models? Starchild-1 and Odyssey explained
  • NVIDIA Cosmos 3: open physical AI world model
  • FLUX 3 from Black Forest Labs: multimodal, video and robotics
  • Runway GWM Worlds 2: an interactive world model
  • ByteDance Seedance as a world model
  • Tencent HY-World 2

Scores, costs and rankings come from Odyssey's October 8, 2026 launch post and the Physics-IQ Verified leaderboard graphic dated October 7, 2026, and may change. Many figures are company-run. We have not run the model. The leaderboard graphic itself is not reproduced here; the table is our own transcription of its text.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 6, 2026

Reka Rho-1: A 19B Model That Reasons, Generates Video and Outputs Robot Actions

Reka AI released a research preview of Rho-1 on October 5, 2026: a 19-billion-parameter omni-reasoning model that treats text, images, video and robot actions as tokens in one shared context, with no tool calls or external models. It was trained on 320 H100 GPUs in about three months. There are no published benchmarks yet. Here is what the design means and what to check.

Aug 26, 2026

Skild AI S1: Robot Tasks From One Video Demonstration

Skild AI's S1 model (August 25, 2026) claims the first long-horizon robotics in-context learner: show a video of pour-over coffee or kit assembly and the robot executes dozens of steps it never saw in pre-training. explainx.ai maps the 7× unseen-task gain Skild reports, how it differs from language-prompted robots, and what NVIDIA physical-AI stacks imply for deployment.

Aug 11, 2026

DYNA-2 World-Action Model and the Robotics Scaling Law Claim

Dyna Robotics unveiled DYNA-2 on August 10, 2026, a "world-action model" pre-trained on over 1,000,000 hours of egocentric human video with no robot data at all. explainx.ai unpacks what a world-action model is versus a VLA, what a scaling law actually claims, and what the published exponents do and don't prove.