Embodied Intelligence Observer

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

ResearchAI 88

Source: arXiv cs.AIPublish time unverified

arXiv:2605.27284v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with coarse goal-level language, leaving execution-critical details such as active arm, approach direction, and contact region unspecified.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies | Embodied Intelligence Observer