Orkas Orkas
Home Blog Research
Research

Vidu S2 vs PixVerse R2: Two Routes to Real-Time Video

Both generate video while you watch and take new instructions mid-stream. Read side by side, the S2 paper and the R2 report agree on the skeleton, split on five design choices — training, drift, memory, speed and control — and disclose very different amounts.

Offline video generation denoises every frame of a clip together; real-time models such as Vidu S2 and PixVerse R2 generate block by block, play each block as soon as it is ready, and apply new input to the next block
Offline models denoise a whole clip at once; real-time models such as Vidu S2 and PixVerse R2 generate block by block and play each block as soon as it is ready. Diagram by Orkas.

Two real-time video models shipped a week apart this month. ShengShu's Vidu S2 went live on September 15, and PixVerse announced PixVerse R2 on September 22. S2 is built around live digital characters and restyling a video stream while it plays. R2 calls itself a real-time world model: a scene you move through with WASD while prompts, references and audio change what happens next.

Both teams published how they built it — S2 as an arXiv paper, R2 as a technical report on the PixVerse site. We read them side by side. They agree on the skeleton, split on five design choices, and disclose very different amounts. This post covers all three and deliberately stops short of naming a winner: the two have never been measured on the same test.

If you make video Real-time is for live experiences. Finished videos still need a workflow. Orkas' VideoStudio agent plans the cut, generates shots and edits them together — with Orkas-managed video models or your own API key.
Download Orkas — free

What real-time changes

Most video models are offline. Every frame of a clip is denoised together over dozens of steps, and nothing is visible until the whole clip is done. Whatever you want has to be in the prompt before generation starts.

A real-time model turns that around. It cuts the video into blocks — a few frames plus the matching slice of audio — and generates them in order. Each block gets only a few denoising steps and plays as soon as it is ready. A block can see what came before it but never what comes after; that is what block-causal means. A new instruction lands in the next block.

That shape creates three problems, and nearly every design choice below answers one of them:

  • Errors compound. Every block is conditioned on blocks the model generated itself, so a small mistake is inherited by everything after it.
  • History has to stay bounded. A stream can run indefinitely; keeping every past block would make each new one more expensive.
  • The time budget is brutal. At 25 frames per second, each frame gets 40 milliseconds.

S2 and R2 land on the same skeleton: a block-causal autoregressive diffusion model that generates video and audio together. The differences are in how they train it, keep it stable, give it memory, make it fast, and let people steer it.

The two systems

Vidu S2 (ShengShu Technology with Tsinghua University) is two models. S2-Avatar is a live digital character at 720p and 25–42 FPS, up from 540p in S1. You can hand it a new reference image mid-stream — a cup, a jacket, a beach — and the character picks it up, puts it on or walks into the scene; it can also dance. S2-Editing restyles an incoming video stream in real time — more than 50 styles, virtual try-on, subject and background replacement — while the source motion stays intact. The paper also explores stereo output for VR headsets.

PixVerse R2 is one model aimed at interactive worlds. Four kinds of input can arrive while it runs — text, multimodal references, audio, and actions such as WASD — and each one updates the state of the world, not just the current frame. PixVerse's game engine now runs on R2; the public demo focuses on movement and prompts.

Decision 1: how the model is trained

Real-time models are not trained from nothing. The usual starting point is a strong offline generative prior, which is then converted to run causally in a few steps. This is where the two reports disagree most openly.

S2 runs a relay. A bidirectional audio–video model is pretrained and then tuned with Diffusion-DPO for fidelity, expression, motion and audio–visual sync. Its attention is switched to block-causal and trained on a mix of clean and noisy history (decision 2). Self-Replay Forcing then distills it to a few steps while it trains on its own outputs, and a final streaming preference pass (Streaming NFT) tunes the causal model. Avatar and editing are trained as separate models.

R2 keeps one foundation. Omni Causal AR is a causal model pretrained continuously on short clips, long videos, multimodal data and interaction trajectories. Real-Time Acceleration then distills that same model for live use: the student is initialized from it and the teacher is built on it. The report's phrase is "accelerate, do not relearn."

Vidu S2PixVerse R2
Starting pointBidirectional model tuned with Diffusion-DPOContinuously pretrained causal model
Route to real timeBlock-causal adaptation → Self-Replay Forcing → streaming preference tuningDistill the same model directly
ModelsSeparate avatar and editing modelsOne model for every input type

R2's report draws the common approach as a five-stage relay and argues that every handoff loses some capability. Read plainly, S2's pipeline has that multi-stage shape, with its main improvement in the last stage. Neither team has published an experiment that compares the two routes, so which one scales better is still an open question.

Decision 2: keeping the stream from drifting

Drift is the signature failure of streaming video: colour creeps, a face slowly turns into someone else, and eventually the frame breaks. It happens because training shows the model clean history, while at inference it only ever sees its own imperfect output.

Both teams start from the same fix. They mix Teacher Forcing, which conditions on clean ground-truth history and protects quality, with Diffusion Forcing, which conditions on history noised to a random level. Noise erases fine detail but keeps layout and motion, so the model learns to lean on structure rather than trust every pixel of the past.

Noised real history is still not the model's own mistakes, so each team adds a second layer.

S2: Self-Replay Forcing. The model first rolls out a long stretch exactly as it would at inference, with no gradients kept. A window of that trajectory is re-noised block by block and replayed in one gradient-enabled causal pass, trained with a DMD distillation loss plus a perceptual loss. Because the replayed blocks sit in one computation graph, gradients cross block boundaries — the model learns how one block shapes the next — without back-propagating through the original rollout.

R2: Error Bank. Representative failure states from generation are stored and replayed during training alongside normal history, so the model learns to recover after a deviation has already entered the world. In PixVerse's internal stage evaluation, a long-horizon brightness-drift metric fell from 0.201 to 0.129, a 35.8% reduction; 20 of 29 long sequences improved, and spurious motion dropped in all five no-motion samples.

The two are complementary rather than competing. Self-Replay Forcing trains on whatever the current model gets wrong; Error Bank drills the failures worth remembering.

Decision 3: memory that stays bounded

Both keep a few opening blocks permanently as an anchor (a sink), keep a sliding window of recent blocks and drop the rest, so the cost of a new block does not grow with the length of the stream. Both also keep positional coordinates inside the range seen in training — S1 calls it RoPE repositioning, R2 calls it relative temporal RoPE — so a long session never pushes positions out of distribution.

S2 builds on TwinCache from S1, where each past block is cached twice: once noisy, once clean. Intermediate denoising steps read the noisy copy, which carries coarse motion and acts like a low-pass filter against accumulating artifacts; the final step reads the clean copy to restore detail. S2 splits this across its two stages: the backbone reads a high-noise cache, the Refiner a low-noise, high-resolution one.

R2 separates memory by timescale: Sink Memory for identity, environment, style and world rules; Rolling History for recent motion, pose and camera; and an Object KV Cache that compresses object-level state that will still matter later.

The split mirrors the products. A digital character has to stay the same person. A world also has to remember what happened in it — the object you put down, the choice you made.

Decision 4: where the speed comes from

Both use sparse attention and few-step distillation. They put their effort in different places.

S2 leans on systems engineering, and documents it. Attention is chosen per layer from SageAttention, SpargeAttention and sparse-linear attention, with the most aggressive approximations on the least sensitive layers. Linear layers run as per-block W8A8 matrix multiplications. Adjacent operators are fused into Triton/CUDA kernels and replayed with CUDA Graphs. Multi-GPU runs use Ulysses context parallelism with quantized communication, and in the editing pipeline the VAE encoder, backbone, Refiner and decoder share GPUs on a common timeline. Resolution comes from a low-resolution backbone plus a one-step latent Refiner up to 720p.

R2 leans on the model. Block-sparse attention is learned during training and reaches more than 90% sparsity. Distillation follows Decoupled DMD — following the control signal and matching the teacher are optimized as separate objectives — plus an adversarial term from DMD2 to anchor realism. Resolution follows a pyramid: one or two low-resolution stages set layout, motion and camera, and a final high-resolution stage adds texture. The report does not state R2's output resolution or frame rate.

Decision 5: how people steer it

S2 wraps the model in a VLM agent. The agent classifies each reference image as a held object, a background or clothing, then writes a prompt for every segment covering identity, expression, gaze, pose, action and held objects, keeping anything the user did not ask to change. After generation it reviews frames in order, judges whether the action finished, half-finished or went wrong, and writes the next prompt accordingly. For putting on or taking off accessories, prompts state both the motion and the end state, so a hat put back on stays on. In editing mode, frame-aligned attention lets each output frame read only the source frame at the same moment, so motion and timing match the input exactly.

R2 builds one input interface into the model. Text, references, audio, actions and agent-generated controls all enter the same running world. Block length follows the active control, up to a cap: a key press gets short blocks for fast response, while a whole event or a stretch of audio gets longer ones to stay coherent. On top of the model, PixVerse's game engine adds an agent layer that keeps game-rule state and the generated scene in sync.

What each report discloses

Before comparing numbers, compare what was published.

Vidu S2PixVerse R2
FormatarXiv paperTechnical post on the PixVerse site
Resolution and frame rate720p, 25–42 FPSNot stated
Parameters and end-to-end latencyNot statedNot stated
Public benchmarksOne avatar benchmark, four editing benchmarksNone; internal evaluation only
WeightsNot released; API availableNot released

On StreamAV-Bench, S2 ranks first on all nine reported metrics across 14 systems in its own evaluation: audio–visual alignment of 0.353 against a best alternative of 0.272, and a synchronization error of 0.617 against 0.648. Some margins sit in the third decimal place — subject consistency is 0.998 against 0.997. R2's published numbers are the 35.8% drift reduction and the 90%-plus attention sparsity, which the report says preserves four internal quality dimensions without giving scores.

There is no head-to-head. The S2 paper compares against the previous PixVerse R1, and R2 was released after it.

What this means if you make video with agents

We build a desktop app where agents do the work, and video is one of the jobs people hand them. Two things from these reports carry over.

First, the line between live and finished video is getting sharper. Real-time models are built for experiences that keep responding — characters, games, a stream you restyle as it plays. Most creator work still ends in a file. In Orkas, VideoStudio turns footage and a brief into a reviewable cut, and when a shot has to be generated it uses offline models: Orkas-managed video generation, or your own key for Seedance 2.0, Hailuo 2.3, Vidu Q3 Pro, Kling 3.0 Turbo, Veo 3.1 or Runway Gen-4.5. Neither S2 nor R2 runs inside Orkas today.

Second, S2's control layer is an agent loop: write the prompt, generate, look at the frames, decide what to do next. That is the same shape as any agent working with a generation model, real-time or not — and much of the result depends on that loop, not only on the weights.

What we take from it

The skeleton has settled: block-causal autoregressive diffusion, video and audio generated together, clean-plus-noisy history, a sink and a window, DMD-family distillation, sparse attention, low resolution first. What separates S2 and R2 is where they spend their effort — a relay of targeted fixes versus one foundation distilled once, on-policy replay versus a bank of failures, systems engineering versus learned sparsity.

Both are worth reading in full: the Vidu S2 paper for its training and serving detail, and the PixVerse R2 report for its argument about scaling a real-time model without relearning it.

For finished videos rather than live streams, the VideoStudio agentic video editing page shows how Orkas takes raw footage to a reviewable cut.