🤖 WSA-Base — World-Spatial-Action Robot Control

A 3D-centric embodied foundation model for generalizable robot control (paper · code · model).

Give it a head-camera view + wrist-camera view of a tabletop robot scene and a task instruction; WSA predicts a short chunk of 7-DoF end-effector actions and can decode its world-model prediction of the next frame. Weights: zaleni/WSA-Base-LIBERO (3B, Qwen3-VL-2B backbone, LIBERO adapter).

Examples
Head camera (agentview) Wrist camera (eye-in-hand) Task instruction