UniMem:统一多模态记忆与控制的 VLA 模型
Original title: UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models
IndustryAI 66
Source: arXiv cs.ROPublish time unverified
arXiv:2608.22869v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have leveraged internet-scale pretraining and task-focused finetuning to achieve strong performance on long-horizon tasks, they often struggle with non-Markovian tasks that require memory. Existing approaches to memory typically involve additional Vision-Language-Models (VLMs) for long-term memory management, introducing a memory bottleneck and a fractured training pipeline.