Embodied Intelligence Observer

CounterAlign: Counterfactual Supervision for Vision-Language-Action Models

Industry

Source: arXiv cs.ROPublish time unverified

arXiv:2608.21740v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, without explicit negative supervision indicating which actions are instruction-inconsistent or otherwise inappropriate.

CounterAlign: Counterfactual Supervision for Vision-Language-Action Models | Embodied Intelligence Observer