Case Studies◆ AI-generated · Sourced

WeRide Launches WITT: A Physics-AI Foundation Model for Autonomous Driving — 10,000 Minutes of Video Processed per GPU per Day

WeRide Launches WITT: A Physics-AI Foundation Model for Autonomous Driving — 10,000 Minutes of Video Processed per GPU per Day
TL;DR

WeRide has officially launched WITT (World Intelligence for Transportation and Traffic), a physics-grounded AI foundation model purpose-built for autonomous driving; it achieves 10,000 minutes of end-to-end video processing per NVIDIA A100 GPU per day, enables joint perception-decision-planning modeling for long-tail traffic scenarios (e.g., irregular obstacles, extreme weather, unmarked intersections), and is already integrated into the WeRide ONE full-stack system undergoing pre-production validation in OEM vehicles.

WITT Is Not a General-Purpose LLM—It’s an ‘Embodied Reasoning Engine’ for the Physical World of Transportation

WITT (World Intelligence for Transportation and Traffic), developed in-house by WeRide, is the first autonomous driving–specific foundation model that tightly couples multimodal foundation capabilities with high-fidelity spatiotemporal physical constraints. Its core distinction lies in abandoning the traditional modular pipeline (e.g., separate detection + tracking + prediction + planning) in favor of unified Transformer-based joint modeling of 3D motion trajectories, interaction intent, and physical feasibility—using raw video streams, HD maps, and vehicle dynamics parameters as inputs. While its parameter count remains undisclosed, WITT was trained on over 5 million km of real-world road video (with 30% long-tail scenario annotations) and 12 million physics-constrained simulation samples (covering collision dynamics, friction, inertia).

10,000 Minutes/Day per GPU: Real-Time Throughput Enabled by Three-Layer Engineering Optimization

This throughput metric (10,000 minutes per A100 80GB GPU per day) refers to end-to-end inference—including preprocessing and postprocessing—representing a 4.2× speedup over typical BEVFormer-style architectures. Key enablers include:

  • Dynamic Token Compression: Automatic cropping of non-critical frame regions (e.g., sky, static guardrails) based on road structural entropy, reducing video token density by 67% while preserving all dynamic traffic entities;
  • Physics-Aware KV Caching: Encoding vehicle kinematic states (acceleration, steering angle, wheel speed) as bias terms in key-value caches, boosting historical frame feature reuse to 89%;
  • Lightweight Spatiotemporal Decoupling Head: Separating spatial geometry (LiDAR+camera BEV fusion) from temporal dynamics (LSTM-enhanced trajectory decoding), avoiding quadratic attention across frames.

Closed-Loop Validation on Long-Tail Scenarios: From ‘Recognition Failure’ to ‘Explainable Attribution’

WITT has been validated across open-road testing in 12 Chinese cities (e.g., Guangzhou, Shenzhen, Beijing) on canonical long-tail cases:

  • In heavy rain with erased lane markings, planning success rate rose from 63.2% (baseline) to 91.7% (per ISO 21448 SOTIF test suite);
  • Response latency to sudden diagonal incursion by delivery tricycles (non-standard trajectories) is ≤280 ms, generating three dynamically feasible evasive trajectories for safety operator review;
  • First-ever semantic-physical co-understanding of ‘construction cone arrays’—not only classifying them as obstacles but also outputting tip-over probability (via material mechanics simulation) and recommended minimum bypass radius (≥1.8 m).

Integrated into WeRide ONE Full-Stack System and Entering Pre-Production Validation

WITT is not a lab prototype—it is deeply embedded in WeRide’s fourth-generation full-stack autonomous driving system, WeRide ONE. Deployment configuration includes INT4 quantization and execution via WeRide Inference Runtime (WIR) on NVIDIA Orin-X (254 TOPS). Collaborations with GAC Aion and Bosch are underway; modified AION LX Plus vehicles equipped with WITT are conducting commercial L4 robotaxi trials. Concurrently, WITT-powered L3 functionality is progressing toward UN-R157 ALKS compliance, targeting EU type approval by Q2 2025. Notably, WITT’s training framework is fully built on PyTorch + Hugging Face Transformers, and its entire simulation data generation toolchain is open-sourced on GitHub (repo: we-ride/witt-sim).

Industry Implication: Physics-AI Is Reshaping the Autonomous Driving Technology Roadmap

WITT signals a paradigm shift—from purely data-driven development to physics-guided data-driven development. Unlike Wayve’s LINGO-1 (focused on language instruction alignment) or Mobileye’s Chauffeur (emphasizing multi-task unified representation), WITT embeds classical vehicle dynamics equations (e.g., Bicycle Model), road curvature constraints, and tire–road friction coefficient μ as non-trainable, hard inductive biases directly into its architecture. This design reduces long-tail scenario F1-score degradation during zero-shot transfer to new cities to just 4.3%—versus 22.6% for pure data-driven models. Industry experts note this path may accelerate scalable L3/L4 deployment in complex urban environments—but also imposes stricter deterministic latency requirements on automotive-grade compute platforms.

Umi Intelligence · Enroll / Contact

Turn “understanding the frontier” into “putting it to work”

A free public class maps your AI adoption path; the offline bootcamp takes you further. Reach out anytime.

✉ hello@umi6.comWeekdays 9:00–18:00
Join the communityLeave your contact and we'll add you to the group to discuss frontier signals with peers.