具身智能观察

CounterAlign: Counterfactual Supervision for Vision-Language-Action Models

产业动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.21740v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, without explicit negative supervision indicating which actions are instruction-inconsistent or otherwise inappropriate.

CounterAlign: Counterfactual Supervision for Vision-Language-Action Models | 具身智能观察