Beyond VLAs: How World Action Models Reshape Robot Manipulation
NVIDIA's Cosmos-based World Action Models aim to generalize robot manipulation beyond training demonstrations
NVIDIA's developer blog introduces World Action Models (WAMs) as an alternative to Vision-Language-Action (VLA) models for robot manipulation, arguing that grounding policies in world physics enables better generalization beyond training demonstrations. The key insight is that robots need to understand underlying physics rather than mimicking demonstrations to handle novel object shapes, positions, or lighting. This is relevant to the physical AI trend but the content is truncated and lacks concrete benchmark results.