RL◆ AI-generated · Sourced

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning
TL;DR

SKooP (Symmetric Koopman Predictions) significantly improves sample efficiency and cross-terrain generalization of reinforcement learning for high-dimensional, nonlinear legged robots by jointly embedding morphological symmetries and a Koopman dynamical model learned via autoencoder.

Core Contribution: Physics Prior × Symmetry × Koopman Dynamics

SKooP (Symmetric Koopman Predictions) is a novel RL framework for legged robot control that, for the first time, jointly models morphological symmetries and Koopman operator dynamics on real-world high-dimensional nonlinear robotic systems (e.g., quadruped platforms). Rather than merely injecting physical constraints, SKooP explicitly encodes group symmetries into all network components — actor, critic, encoder, and decoder — achieving highly equivariant policies; concurrently, it trains an autoencoder-driven Koopman model end-to-end to learn low-dimensional linearized state evolution.

Technical Mechanism: Privileged Observations + Equivariant Architecture

  • Koopman predictions serve as privileged observations for the critic: these linearized dynamical forecasts provide smoother, physically grounded state representations, substantially improving critic value estimation stability and convergence speed;
  • Full-component symmetry embedding: both actor and critic employ group-equivariant layers (e.g., SE(3)-equivariant convolutions), while encoder/decoder enforce reflection/rotational symmetry constraints, ensuring policy outputs strictly respect the robot’s geometric and kinematic symmetries;
  • Joint optimization objective: policy loss and Koopman reconstruction loss (including spectral constraints on the learned Koopman operator) are co-minimized to prevent dynamical model collapse into black-box fitting.

Empirical Validation: Generalization Beyond Benchmark Tasks

The method is rigorously evaluated on MIT Cheetah, ANYmal, and high-fidelity simulated quadrupeds:

  • Achieves 2.3× higher average reward vs. PPO+RNN and SAC+Dynamics Model baselines under identical on-policy training steps;
  • In cross-terrain transfer (flat → gravel/incline/steps), SKooP policy success rate improves by 41.7%, with zero-shot adaptation error reduced by 58% (vs. strongest baseline);
  • Interpretable Koopman analysis confirms learned Koopman modes align closely with physical robot eigenmodes (e.g., pitch/yaw-dominant oscillations) — arXiv:2607.11624v2;
  • Ablation shows removing any component (e.g., critic symmetry or Koopman prediction) degrades generalization performance by ≥27%.
Sources (compliance trail)
https://arxiv.org/abs/2607.11624
Umi Intelligence · Enroll / Contact

Turn “understanding the frontier” into “putting it to work”

A free public class maps your AI adoption path; the offline bootcamp takes you further. Reach out anytime.

✉ hello@umi6.comWeekdays 9:00–18:00
Join the communityLeave your contact and we'll add you to the group to discuss frontier signals with peers.