Step Inside Worlds ThatGenerate Around You — In Real Time
GWM Worlds 2 turns high-fidelity video and audio generation into interactive simulation. You define the environment, subjects, visual style, physical rules and ambience — then steer the world with text actions addressed to any subject or the scene itself, alongside continuous camera motion. Sessions have no preset length.
See GWM Worlds 2 Generated Live
Every clip below was generated live at 24 fps, with a user steering the world through text actions and continuous camera motion.
Playing the survivor: a desert world played in first person — the survivor moves, jabs the spear, drinks from a canteen and breaks into a run. Directing the scene: prompts addressed to the campfire and the weather flare the fire and fade the sunset into a starry night. Both at once: a player crosses the desert while a director reshapes the scene around them.
More Worlds, Live
WorldPrompt: The Input Format for Worlds
Rich control over a world requires an input representation that generalizes. WorldPrompt splits a world into two kinds of state: what persists and what changes over time.

Scene. An urban wet pavement lined with brick buildings and stone trim. Reflections are visible on the wet ground, scattered with yellow autumn leaves.
Subject. A street sweeper in an orange high-visibility vest holds a broom and dustpan shovel, and speaks with a male voice. A woman in a green trench coat walks down the wet pavement.
Law. Gravity behaves like Earth, leaves lie flat on the wet pavement, and the camera follows the woman in third person.
Play the clip — the world follows the timestamped event stream: sweeping, pointing, walking, and spoken lines.
Finetuned on WorldPrompt — Following the Structure Correctly
Why GWM Worlds 2 Is a Research Breakthrough
Building on GWM Worlds, Runway post-trains its foundational audio-video model into a real-time autoregressive world model — continuous video and audio, indefinitely.
GWM Worlds 2 generates 720p video at 24 fps and audio at 48,000 Hz as you play. Unlike the bidirectional base model, it is not restricted to a fixed duration and can generate indefinitely.
A text action can be a sword slash, a spoken reply, a room flooding with water, a dust storm rolling in — or anything else you can put into words. Address any subject, or the scene itself.
Audio is generated natively alongside video at 48,000 Hz. Characters hold conversations and control the voice, tone and language of speech — with lip movement and delivery generated to match.
First-person and third-person navigation — walking, driving, riding — with the camera and the subject controlled independently. Move through a world however you like.
Agents can steer subjects to accomplish goals, making GWM Worlds 2 a simulated environment for evaluation and training. Multiple users can control different subjects in shared worlds.
Prefill generation with an existing video instead of a first frame — a generated clip, a shot from your edit — and continue playing from it with environment and audio kept consistent.
Built to Be Steered, Played and Directed
The same WorldPrompt structure applies live — the world responds as you play. Recurring actions can be bound to keys for fast play.
First-person and third-person navigation — walking, driving and riding through a world. The camera and the subject can be controlled independently.
Direct subjects and the scene to take arbitrary actions: running, climbing and leaping; switching a lamp on or changing its color; a theater stage erupting in fire.
Characters hold conversations and control the voice, tone and language of the speech — NPC–player dialogue or a vlogger talking to camera, with lip movement matching the delivery.
Three Ways to Use a World Model
The model takes a WorldPrompt and outputs video and audio — in whatever rhythm your workflow demands.
Author the entire timestamped event stream once at the start — with or without an LLM — and the model renders the full video and audio session with it.
Generation runs until a decision point. You pick an action, an LLM casts it into the event stream, and the model continues until the next choice.
The world generates continuously and reacts to actions with minimal latency — the most challenging scenario, built for games and live interactive experiences.
Instead of typing prompts on the fly, the demo binds key and mouse inputs to premade prompts — the W key might move the character forward, left-click might throw a ball. Mouse movement steers the camera.
Prefill generation with an existing video — here, a generated eight-second clip — and keep playing from it. The model keeps the environment and audio consistent with the input video.
Agents can control subjects and the environment — steering an adventurer through parries and sword slashes while directing castle lighting, or riding a jet ski across waves.
A robot was instructed to reach the red then blue flag — and succeeded on its first try both times. GWM Worlds 2 works as a simulated environment for evaluating and training agents.
Different users control different subjects. Roles — player 1, player 2, director — each carry their own actions addressed to their own subjects, with video and audio broadcast via LiveKit.
Go from idea to world within seconds. Type a simple prompt like "third person perspective dirt bike in a snowy landscape" and an LLM assistant generates your first frame, drafts the genesis prompt and binds subject actions to keys.




Foundations for Entertainment, Robots & Research
Because the world continues from each new input instead of following a fixed script, GWM Worlds 2 opens entirely new categories.
Games, visual novels and interactive films where the world reacts to every decision — with no pre-baked clips and no preset session length.
Speaking, reacting characters with generated voices and lip-synced delivery — NPCs, vloggers, virtual companions that hold real conversations.
A simulated environment for evaluating and training agents in diverse, controllable worlds — with goal-driven control over subjects and the scene.
Author full scenes ahead of time and render complete video-plus-audio sessions, or scout shots by steering a camera through a world live.
Generative environments and interfaces that adapt to input in real time — new kinds of design tools and interactive products built on worlds.
A research preview pushing real-time video generation forward — the same curve that took offline video from rough clips to production footage is now playing out live.
Under the Hood
GWM Worlds 2 is an autoregressive diffusion video and audio model.
Autoregressive Diffusion
720p video at 24 fps and audio at 48,000 Hz, generating indefinitely with no fixed duration.
Three Contexts per Step
Global context (genesis prompt + first frame), the current frame's inputs (camera + text actions), and past frames cached in a sliding window.
Causal Caching Decoders
Video and audio decoders are causal and run with a cache for faster decoding.
Sliding KV-Cache Window
Each frame's video, audio, text and camera tokens attend to the global tokens and past frames; older frames are evicted.
Frequently Asked Questions
Common questions about GWM Worlds 2 and how to get started.
Research preview. GWM Worlds 2 is an early point on the real-time generation curve. Constrained by fidelity-versus-speed tradeoffs, sensitive to fast camera rotations, and limited in long-term memory — exactly the constraints the Runway research team says will be solved with continued work.
Ready to Step Inside aGenerative World?
GWM Worlds 2 turns video and audio generation into real-time interactive simulation — worlds that keep going as long as you do.