Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
产业动态AI 67
来源:arXiv cs.RO发布时间待核实
arXiv:2608.23478v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction.