NVIDIA researchers published Long-WAM on October 7, 2026, and it reached the number two spot on Hugging Face's Papers of the Day on October 8. The paper asks a simple question for robot control: how much video history should a policy see, and what does it take to make that history useful?
The answer has two parts. Longer context helps a lot, going from 63.3% to 78.7% success on the RoboCasa GR-1 benchmark when context grows to 19.2 seconds. But it only helps when the video backbone was pretrained autoregressively. A bidirectionally pretrained backbone shows no net gain. The authors put it this way: "access to history is not the same as using it."
This post explains what a world-action model is, how Long-WAM is built, what the numbers say, and where the results stop.
TL;DR: Long-WAM in one table
| Question | Answer |
|---|---|
| Who built it? | NVIDIA, with MIT, HKU and UCSD authors. 16 authors including Jim (Linxi) Fan, Song Han and Yukang Chen |
| What is it? | A framework for scaling visual context in causal world-action models under real-time limits |
| Headline result | RoboCasa GR-1: 63.3% to 78.7% when context goes from 0.0 to 19.2 seconds |
| Other results | LIBERO-Long 99.5%, RoboTwin 2.0 94.4% average, DOMINO 34.9% |
| Speed | 107.4 ms per action chunk on RTX 5090, including future prediction |
| Real robot | Unitree G1 dynamic cup stacking 19 of 20 trials. Both baselines 0 of 20 |
| Where to read it | arXiv 2610.10528, Hugging Face Papers, the NVlabs project page |
What is a world-action model?
A world-action model (WAM) is an embodied foundation model that predicts the future and picks actions in one network. A standard vision-language-action policy maps images and an instruction to an action. A WAM also models how the scene will evolve. It learns "what happens next" from large amounts of video, then uses that physical knowledge to choose motions.
This is the same family of ideas that explainx.ai covered in what are world models and in NVIDIA Cosmos 3. It is also close to Reka Rho-1, which emits robot actions from a shared network. WAMs are a fast-moving area. Hugging Face's recommendation bot listed seven similar October papers next to Long-WAM, including Streaming-WAM, DualWAM, Rolling-WAM and OpenWAM.
Why does context length matter for robots?

Real-time control demands enough visual history to infer motion and task progress. One frame cannot show whether a cup is moving or how far a task has advanced. But processing a long history takes time, and a slow policy reacts late. The abstract states this tension directly: processing history "can delay action."
Think of a robot catching an object on a moving conveyor. The speed of the object is not in a single image. The task stage, such as "the cup is already stacked, now grasp the next one," is not in a single image either. Memory lets the policy infer both. The cost is latency, and latency causes misses.
Long-WAM attacks both sides. It scales the context, and it rebuilds the runtime so that the longer context does not make the robot too slow.
How does Long-WAM work?
The project page describes a "Remember, imagine, act" loop with four stages.
- Autoregressive video pretraining. The team continues the LongLive-2.0 autoregressive checkpoint into LongLive2.0-Robot. It trains on roughly 10,000 window-equivalent hours of robot and egocentric video from RoVid-X, AgiBot World, EgoDex, EgoVerse and VITRA. The model learns teacher-forced next-chunk prediction on sequences up to 30 seconds, without action labels.
- Causal-to-causal world-action adaptation. A video expert and an action expert are coupled through an asymmetric attention interface, a Mixture of Transformers. Video queries never read action tokens. Action queries read the observed history, the partially denoised future latents and the action chunk. The model predicts the future first and then acts, without decoding pixels.
- Context scaling. The team varies the duration of real observations in the causal prefix and keeps the forecast and action horizons fixed. Each window is trained separately and evaluated at its own length, up to 38.4 seconds on RoboCasa GR-1.
- Real-time infrastructure. Pure asynchronous execution, streaming VAE encoding, NVFP4 quantization, KV reuse, CUDA Graphs and device-specific kernel tuning bring the full model to the RTX 5090, DGX Spark and Jetson AGX Thor.
In the inverse-dynamics mode, the video expert predicts future latents from the observed history. The action expert then denoises an action chunk conditioned on that history and the predicted future. Skipping pixel decoding saves time.
What do the benchmark numbers show?
The headline claim is about context scaling. On RoboCasa GR-1, the paper reports that increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%. The project page's summary says 2.4 to 19.2 seconds raises success from 66.3% to 78.7%. Both ranges end at the same 78.7%. They differ only in the starting point, and a reader should check the paper's table when comparing.
Other reported results on the project page:
| Benchmark | Result |
|---|---|
| RoboCasa GR-1 | 78.7% with 19.2 s of context |
| LIBERO-Long | 99.5% (from 94.5% with 2.4 s of history) |
| RoboTwin 2.0 | 94.4% average, clean plus randomized. 94.2% when run asynchronously |
| DOMINO | 34.9% after dynamic-data fine-tuning, highest among compared methods |
| RoboCasa365, 50 tasks | 54.4% when coupled with GPT-6 Astra as planner. 31.4% without it |
| RoboCasa365, unseen compositions | 35.0% with GPT-6 Astra, up from 6.1% |
The most interesting row is the control. The authors compare three video initializations: LongLive2.0-Robot (autoregressive), LongLive2.0 (autoregressive) and Wan2.2 (bidirectional). Both autoregressive versions turn longer history into higher success. The bidirectional version shows no net gain on GR-1 between 2.4 and 19.2 seconds. The authors conclude that the value of context depends on autoregressive pretraining.
What happened on real robots?

Credit: NVIDIA Long-WAM project page (poster frame).
Real-world tests use a Unitree G1 and a YAM arm. According to the project page:
- Moving conveyor. The G1 grasps objects from a conveyor moving up to 7.5 cm/s. Both baselines, pi 0.5 and Fast-WAM, fail every trial at 6.0 and 7.5 cm/s.
- Dynamic cup stacking. Long-WAM succeeds in 19 of 20 trials. Both baselines succeed in 0 of 20.
- Long horizon on YAM. Tasks lasting over 40 seconds average 81.7% success.
These are small samples with 20 trials per policy and speed. They show a clear gap on dynamic tasks, but they are not a statistical benchmark. The project page compares against pi 0.5 and Fast-WAM and includes a human teleoperation reference.
How fast is it, and can it run on a robot?

Speed is half of the contribution. The authors report 107.4 ms per action chunk on an RTX 5090 with future prediction retained. At a 19.2 second context the latency is higher: the page's interactive explorer lists 341.0 ms per chunk at that length on the RTX 5090. So the "real-time" claim depends on the context setting. A reader should check which context length was used for which latency number.
The optimizations bring speedups of 3.3x on RTX 5090, 4.1x on DGX Spark and 3.2x on Jetson AGX Thor, compared with BF16 eager mode. The Jetson result matters most for robot builders, because it is the module that sits inside the machine. For a view of the desktop-class option, see explainx.ai's DGX Spark price breakdown.
Why does this matter for builders?
First, it is a recipe, not only a score. The recipe says: pretrain an autoregressive video model on lots of unlabeled robot and egocentric video, then keep the causal structure when you add actions. If you work with robot policies, that is a testable design choice.
Second, it shows a general lesson that applies to agents too. Giving a model more history does not help if the model was not trained to use it. The same idea appears in language models, where a bigger window is not the same as better use of the window. explainx.ai's LLM context window explainer covers that parallel.
Third, the pairing with a language-model planner is notable. The project page reports that GPT-6 Astra, acting as a planner that decomposes goals, recovers from failures and corrects actions, lifts the unchanged Long-WAM checkpoint from 31.4% to 54.4% overall on RoboCasa365. The split into a reasoning layer and an executor is a pattern seen in other embodied systems, such as Figure's Helix 02.
What are the limits?
The authors are clear about some limits, and some come from the setup.
- Simulation-heavy. Most headline numbers come from simulated benchmarks. Real-world tests are narrow.
- Context ceiling. At 38.4 seconds, 80.4% of sampled history frames are padding, because training trajectories average 12.1 seconds. The page says the drop at that length "reflects limited effective history, not an established memory limit." The true limit of useful memory is not known.
- Latency grows with context. The 107.4 ms figure and the 341.0 ms figure are not the same setting.
- Self-reported. No independent replication exists yet. The paper is a day or two old.
- Compute needs. Training uses about 10,000 window-equivalent hours of video and a large video backbone. The inference stack depends on NVIDIA-specific optimizations such as NVFP4 and CUDA Graphs.
- Open questions on release. Hugging Face lists the GitHub repository as NVlabs/LongLive. Check the repository for license terms and checkpoint availability before you plan on it.
What people are asking
Is this a language model? No. It is a video-and-action model. A language model such as GPT-6 Astra appears only as an optional planner in one experiment.
Is 78.7% a good score? It is the best number the authors report on RoboCasa GR-1 for their own method. The paper compares it with their own ablations and with baselines on other benchmarks. A reader should look at the paper's full comparison table for GR-1 baselines.
Can I run it at home? The project page says the model is deployed on RTX 5090, DGX Spark and Jetson AGX Thor. Whether weights are public is something to confirm in the repository.
Related reading
- What are world models? Starchild-1, Odyssey and the complete guide
- NVIDIA Cosmos 3: open physical AI world model
- Reka Rho-1: one network for text, video and robot actions
- Odyssey-3 tops Physics-IQ Verified
- Figure Helix 02: collaborative humanoid robots
- NVIDIA MotionBricks and Unitree G1
- LLM context windows explained
Primary sources: arXiv 2610.10528; Hugging Face paper page; NVIDIA Long-WAM project page.
Numbers on this page come from the paper abstract and the NVIDIA project page as of October 9, 2026. They are authors' reports and may change in later paper versions.
