New Research Preview — The Next Runway World Model

Step Inside Worlds ThatGenerate Around You — In Real Time

GWM Worlds 2 turns high-fidelity video and audio generation into interactive simulation. You define the environment, subjects, visual style, physical rules and ambience — then steer the world with text actions addressed to any subject or the scene itself, alongside continuous camera motion. Sessions have no preset length.

720p Video @ 24 FPS
Audio @ 48,000 Hz
Text-Steered Actions
No Fixed Length
720p
Real-Time Video
24 FPS
Live Generation
48 kHz
Generated Audio
Unlimited Sessions
By Runway
Agentic & robotics-ready
Multiplayer-ready

See GWM Worlds 2 Generated Live

Every clip below was generated live at 24 fps, with a user steering the world through text actions and continuous camera motion.

Playing the Survivor
Directing the Scene
Both at Once

Playing the survivor: a desert world played in first person — the survivor moves, jabs the spear, drinks from a canteen and breaks into a run. Directing the scene: prompts addressed to the campfire and the weather flare the fire and fade the sunset into a starry night. Both at once: a player crosses the desert while a director reshapes the scene around them.

More Worlds, Live

Dirt Bike — Snow
Jeep — Forest
Oasis Traveler
Unicorn Ride
Fortress Parkour
Lamp Colors
Green Blob
Theater on Fire
Village Market Chat
Vlogger to Camera

WorldPrompt: The Input Format for Worlds

Rich control over a world requires an input representation that generalizes. WorldPrompt splits a world into two kinds of state: what persists and what changes over time.

Persistent World Context
A genesis prompt describing the environment, layout, materials, lighting and ambient sound
The subjects that can participate in events, with their attributes
The laws — reusable conventions governing behavior, from gravity and collision to character abilities and the camera perspective
A first frame to ground the generation visually
Timestamped Event Stream
Actions describing movement, gestures, object interactions, speech and sound — each a free-form text prompt with start and end timestamps, addressed to a subject or the scene itself
Multiple actions can overlap freely
Camera input — a per-frame stream of viewpoint translation and rotation
Speech is an action like any other, carrying the line to be said
First frame of the example session: a street sweeper on the left and a woman in a green trench coat

Scene. An urban wet pavement lined with brick buildings and stone trim. Reflections are visible on the wet ground, scattered with yellow autumn leaves.

Subject. A street sweeper in an orange high-visibility vest holds a broom and dustpan shovel, and speaks with a male voice. A woman in a green trench coat walks down the wet pavement.

Law. Gravity behaves like Earth, leaves lie flat on the wet pavement, and the camera follows the woman in third person.

Play the clip — the world follows the timestamped event stream: sweeping, pointing, walking, and spoken lines.

Finetuned on WorldPrompt — Following the Structure Correctly

Lakeside Chat
Cartoon Castle

Why GWM Worlds 2 Is a Research Breakthrough

Building on GWM Worlds, Runway post-trains its foundational audio-video model into a real-time autoregressive world model — continuous video and audio, indefinitely.

Real-Time Autoregressive Generation

GWM Worlds 2 generates 720p video at 24 fps and audio at 48,000 Hz as you play. Unlike the bidirectional base model, it is not restricted to a fixed duration and can generate indefinitely.

Continuous 720p @ 24 fps
Sliding KV-cache window
Text Actions on Anything

A text action can be a sword slash, a spoken reply, a room flooding with water, a dust storm rolling in — or anything else you can put into words. Address any subject, or the scene itself.

Free-form text control surface
Overlapping simultaneous actions
Generated Audio & Speech

Audio is generated natively alongside video at 48,000 Hz. Characters hold conversations and control the voice, tone and language of speech — with lip movement and delivery generated to match.

NPC–player dialogue
Lip-synced delivery
Independent Camera Control

First-person and third-person navigation — walking, driving, riding — with the camera and the subject controlled independently. Move through a world however you like.

Continuous camera motion
Walk, drive or ride
Agentic & Multiplayer Ready

Agents can steer subjects to accomplish goals, making GWM Worlds 2 a simulated environment for evaluation and training. Multiple users can control different subjects in shared worlds.

Goal-driven agent control
LiveKit multiplayer roles
Continue From Any Video

Prefill generation with an existing video instead of a first frame — a generated clip, a shot from your edit — and continue playing from it with environment and audio kept consistent.

Prefill with generated video
Consistent world & audio

Built to Be Steered, Played and Directed

The same WorldPrompt structure applies live — the world responds as you play. Recurring actions can be bound to keys for fast play.

Navigation

First-person and third-person navigation — walking, driving and riding through a world. The camera and the subject can be controlled independently.

Actions

Direct subjects and the scene to take arbitrary actions: running, climbing and leaping; switching a lamp on or changing its color; a theater stage erupting in fire.

Speech

Characters hold conversations and control the voice, tone and language of the speech — NPC–player dialogue or a vlogger talking to camera, with lip movement matching the delivery.

Climb & Leap
Switch Object States
Scene Events
Talking to Camera
Ride & Steer
Fantasy Rides
First-Person Travel
Creature Actions

Three Ways to Use a World Model

The model takes a WorldPrompt and outputs video and audio — in whatever rhythm your workflow demands.

1
Ahead of Time

Author the entire timestamped event stream once at the start — with or without an LLM — and the model renders the full video and audio session with it.

FilmmakingAdvertising
2
Turn-Based

Generation runs until a decision point. You pick an action, an LLM casts it into the event stream, and the model continues until the next choice.

Visual novelsInteractive film
3
Real-Time

The world generates continuously and reacts to actions with minimal latency — the most challenging scenario, built for games and live interactive experiences.

GamesInteractive experiences
The Real-Time Demo

Instead of typing prompts on the fly, the demo binds key and mouse inputs to premade prompts — the W key might move the character forward, left-click might throw a ball. Mouse movement steers the camera.

A mage in a snowy pass, played live: key and mouse bindings fire premade action prompts.
Continuing Play From a Video

Prefill generation with an existing video — here, a generated eight-second clip — and keep playing from it. The model keeps the environment and audio consistent with the input video.

Urban war zone — 8s prefilled input video, then played live: firing, strafing, sprinting and diving prone.
Agentic Control

Agents can control subjects and the environment — steering an adventurer through parries and sword slashes while directing castle lighting, or riding a jet ski across waves.

A robot was instructed to reach the red then blue flag — and succeeded on its first try both times. GWM Worlds 2 works as a simulated environment for evaluating and training agents.

Multiplayer Worlds

Different users control different subjects. Roles — player 1, player 2, director — each carry their own actions addressed to their own subjects, with video and audio broadcast via LiveKit.

One user rides the jet ski while another directs the world around them.
World Authoring

Go from idea to world within seconds. Type a simple prompt like "third person perspective dirt bike in a snowy landscape" and an LLM assistant generates your first frame, drafts the genesis prompt and binds subject actions to keys.

Describe the world you want
A first frame is generated
The genesis prompt is drafted
Subject actions are bound to keys

Foundations for Entertainment, Robots & Research

Because the world continues from each new input instead of following a fixed script, GWM Worlds 2 opens entirely new categories.

Interactive Entertainment

Games, visual novels and interactive films where the world reacts to every decision — with no pre-baked clips and no preset session length.

Virtual Characters

Speaking, reacting characters with generated voices and lip-synced delivery — NPCs, vloggers, virtual companions that hold real conversations.

Robotics & Embodied Agents

A simulated environment for evaluating and training agents in diverse, controllable worlds — with goal-driven control over subjects and the scene.

Filmmaking & Advertising

Author full scenes ahead of time and render complete video-plus-audio sessions, or scout shots by steering a camera through a world live.

Generative Design & Interfaces

Generative environments and interfaces that adapt to input in real time — new kinds of design tools and interactive products built on worlds.

Research & Simulation

A research preview pushing real-time video generation forward — the same curve that took offline video from rough clips to production footage is now playing out live.

Under the Hood

GWM Worlds 2 is an autoregressive diffusion video and audio model.

Autoregressive Diffusion

720p video at 24 fps and audio at 48,000 Hz, generating indefinitely with no fixed duration.

Three Contexts per Step

Global context (genesis prompt + first frame), the current frame's inputs (camera + text actions), and past frames cached in a sliding window.

Causal Caching Decoders

Video and audio decoders are causal and run with a cache for faster decoding.

Sliding KV-Cache Window

Each frame's video, audio, text and camera tokens attend to the global tokens and past frames; older frames are evicted.

Frequently Asked Questions

Common questions about GWM Worlds 2 and how to get started.

Research preview. GWM Worlds 2 is an early point on the real-time generation curve. Constrained by fidelity-versus-speed tradeoffs, sensitive to fast camera rotations, and limited in long-term memory — exactly the constraints the Runway research team says will be solved with continued work.

The Next Runway World Model

Ready to Step Inside aGenerative World?

GWM Worlds 2 turns video and audio generation into real-time interactive simulation — worlds that keep going as long as you do.

No credit card required
Free credits included
Generates in real time