Inference◆ AI-generated · Sourced

Alibaba's Wan-Streamer v0.2: A Real-Time Omni-Modal "AI Video Call" at 550ms End-to-End

Alibaba's Wan-Streamer v0.2: A Real-Time Omni-Modal "AI Video Call" at 550ms End-to-End
TL;DR

Tongyi Lab unifies listen-see-speak-perform in a single Transformer, hitting ~550ms end-to-end latency and 640×368@25FPS, using a decoupled Thinker/Performer dual path to keep latency low while raising resolution.

Making AI "listen, watch and respond" like a real person

Per IThome (July 17), Alibaba's Tongyi Lab released Wan-Streamer v0.2 — an end-to-end, full-duplex omni-modal understanding-and-generation model that unifies listening, seeing, speaking and performing in one Transformer. Three headline metrics: ~550ms end-to-end latency (200ms model + 350ms network); output resolution raised from v0.1's 192×336 to 640×368 @ 25FPS with clear micro-expressions; and native real-time understanding and synchronized generation of text, audio and video without external modules. The team says it beats common real-time systems on both latency and feature coverage (video perception, video output, full duplex, end-to-end, sub-1s latency).

Key mechanism: causal timeline + streaming units

Wan-Streamer maps user text/audio/video input and agent output onto a single Causal Timeline and introduces Streaming Units: roughly every 160ms it completes one loop — perceive the current 160ms of A/V input, update shared interaction state and context, generate synchronized speech and video latents, and decode/emit the previous unit's response. The AI doesn't wait for you to finish; within each slice of your speech it runs perceive→understand→generate→decode. This native streaming design underlies its ultra-low latency and full-duplex interaction.

Higher quality without added latency: Thinker/Performer decoupling

v0.1 proved native streaming A/V dialogue feasible but 192×336 allowed only close-ups. v0.2 raises resolution to 640×368 — the AI is no longer a "floating head," showing gaze, posture, natural gestures and even the room. But high-res video latents are compute-heavy; routed through the low-latency path they would blow past the 200ms model-side floor. So v0.2 splits the model into two parallel paths, decoupled on hardware:

  • Thinker (single-GPU fast lane): handles all latency-sensitive work — streaming A/V perception, language/state updates, K/V cache, audio decoding — preserving the 200ms response.
  • Performer (multi-GPU Ulysses parallelism): carries the heavy 640×368 video generation, expanded into a multi-GPU Ulysses-style context-parallel cluster that splits long video sequences across GPUs for parallel denoising.

By letting a single GPU own low latency and multiple GPUs own high resolution, Wan-Streamer v0.2 improves image quality without increasing user-perceived latency.

Sources (compliance trail)
https://www.ithome.com/0/978/125.htm
Umi Intelligence · Enroll / Contact

Turn “understanding the frontier” into “putting it to work”

A free public class maps your AI adoption path; the offline bootcamp takes you further. Reach out anytime.

✉ hello@umi6.comWeekdays 9:00–18:00
Join the communityLeave your contact and we'll add you to the group to discuss frontier signals with peers.