具身智能观察

Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models

产业动态

来源:NVIDIA 技术博客发布时间待核实

Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it... Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it to generate actions from visual observations and language instructions. Large-scale VLM pretraining is a core part of the recipe. See Pi-0 and GR00T N1.

Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | 具身智能观察