如果说,基础模型是具身大脑的“大学”通识课;那么,垂域模型则是机器人在工业领域的“专业课”。 作者丨董子博 编辑丨林觉民 今天要做世界模型,所有人都知道要做预训练,来让具身大脑变得更聪明;也知道要做后训练,让机器人可以在专门的任务下得到锻炼、做得更好。 和不少公司聊过,AI 科技评论发现,谈到中训练(Mid-train)的公司却很少。 而实际上,在具身的实际落地中,却存在一些原有方案不易解决的问题:当预训练的 VLA 模型,没法在具体的场景中发挥全部智能;而后训练却有时间、成本的门槛;在训练技能的环节,也容易出现数采、训练和调试的重复工作——VLA 落地要真正提效,中训练或许才是破局的良方。 AI科技评论独家获悉,乐聚机器人将推出面向工业场景构建的“垂域具身智能模型”——KUAVO VLA。据了解,该模型以通用VLA 基础模型为能力起点,使用 600+ 小时KUAVO 同构型真机数据进行了中训练,将本体适配、基础操作和工业场景的共性能力沉淀到模型中。 乐聚机器人 KUAVO VLA 赋能下,在不少场景中可以实现大约一倍的提效,在某些任务中甚至可以做到100% 的成功率。 模型中训练,会…

商汤董事长兼首席执行官徐立表示:“商汤构建的‘一套模型、一座 Token 工厂、一套智能体管控系统’核心能力体系,用一套模型实现持续提升智能上限的路径;持续用基础设施和优化能力降低Token 成本,提高成本定价能力;以智能体深度融入企业与个人工作流,形成结果交付能力。这套体系赋能商汤超越单纯价格竞争,凭借差异化的技术与服务,开创可持续商业价值,同时扩大用户基础,带动经常性收入增长,为集团业务的高质量发展奠定坚实基础。” 一套模型统一多模态架构,持续提升智能上限 模型展现国际领先水平 2026 年,商汤密集发布多款模型,覆盖原生多模态统一、视觉感知与推理、可控生成、长程智能体、空间智能与物理世界模型等方向,包括 NEO-Unify 架构、SenseNova U1、U1.5、U1.5-Lite、SenseNova Vision、办公场景的多模态智能体SenseNova 6.8-Flash-Lite、内容创作模型SenseNova U1-Pro、空间智能模型SenseNova SI-8B,以及与大晓机器人联合开发的 Kairos 3.1世界模型。上述模型在多项核心基准上对标并持平或超越 Ge…

arXiv:2608.18234v2 Announce Type: replace Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy t…
arXiv:2606.28455v3 Announce Type: replace Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics. We introduce a controlled diagnostic protocol for studying event-conditioned latent physical structure in passive object-state world models. The protocol separates three questions: whether event-regime information is readable, whether event context changes the relative empha…
arXiv:2606.01027v2 Announce Type: replace Abstract: Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present $\tau_0$-World Model ($\tau_0$-WM), a unified video-action world model that integrates policy learning, video prediction, and action evaluation within a single future-predictive framework. Built on a shared video diffusion backbone, $\tau_0$-WM provides two complementary interfac…
arXiv:2608.22294v1 Announce Type: new Abstract: World models for physical interaction are typically trained to predict future observations or latent features; however, a planning-oriented model must answer a fundamentally different question: whether a candidate action produces a task-consistent future while preserving essential relations.Monolithic state representations obscure the underlying entities, while standard instance-level object slots merely identify \emph{what} is present without spec…
arXiv:2608.22278v1 Announce Type: new Abstract: Vision-based whole-body loco-manipulation on humanoid robots is challenging due to partial observability, contact-rich dynamics, and the difficulty of learning long-horizon behaviors from high-dimensional visual inputs. We present \href{https://github.com/DreamMimic/DreamMimic}{DreamMimic}, a framework that distills privileged teacher policies into vision-based humanoid controllers via world-model-assisted distillation. Instead of using a Dreamer-s…
arXiv:2608.22067v1 Announce Type: new Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed.…
arXiv:2608.21414v1 Announce Type: new Abstract: Autonomous driving risk identification aims to determine which observed object is likely to become safety-critical to the ego vehicle. Existing approaches typically predict scene-level accidents, infer risk objects indirectly from ego behavior, or apply geometric checks after trajectory forecasting, without directly using predicted ego--object relations for risk-source localization. We propose RiskWorld, an object-centric latent world model that id…
AI 生成的大多数世界“中看不中用”,根本原因是 AI 少了一本“底账”:一份记住世界里有什么、每次行动改变了什么、并且能被检查和回滚的状态记录。 作者丨刘紫东 编辑丨幸丽娟 李飞飞押注的空间智能,被视为世界模型的主流路线之一。 这条路线的基本构想是,让 AI 从识别二维画面,走向理解、推理和生成可进入、可操作、可持续编辑的三维世界。 但真正的考验在于:一个“站得住”的世界,是怎样构建出来的? 要回答这个问题,不妨先回到一个更基础的追问:当 AI 的交付形态从“画面”变成“三维世界”,互动方式从“观看”变成“进入和操作”,到底意味着什么? 它意味着 AI 交付的内容开始拥有时间属性:对象需要在多轮操作中保持身份,规则要承受真实行动,局部修改也要与已经成立的关系相容;系统还得从每一次行动的后果中,修正自己对世界的理解。 李飞飞将这组能力概括为 Renderer、Simulator 与 Planner——通俗地说,一个负责把世界“画出来”,一个负责记住“世界现在是什么样、行动之后会变成什么样”,一个负责决定“下一步做什么”;三者之间的连接,尤其是模拟与规划共享的状态层——也就是那本“底账”…

8 月 19 日下午,世界机器人大会分论坛“具身智能大模型的进化之路:从技术分级到产业落地”在北人亦创国际会展中心举行。生数科技创始人兼首席科学家、ACM/IEEE/AAAI Fellow 朱军发表主旨演讲,发布团队最新研究成果,并系统阐述通用世界模型的五级发展路线。 朱军表示:“从大模型发展的视角来看,我们希望构建的,不只是服务于某个任务或场景的专用模型,而是一个能够理解世界、预测未来并采取行动的通用基座模型。” 从第一性原理出发,定义通用世界模型 世界模型已经延伸到视频生成、环境模拟、机器人决策和动作控制等方向,但行业对“什么样的世界模型才称得上通用”仍缺少共识。 人在学习骑自行车或开车时,会从最初的不协调逐渐变得稳定、准确。其重要原因,是大脑在与外部世界持续互动的过程中,形成了一个关于世界如何运行的“内部模型”。 通用世界模型同样需要三项彼此关联的能力: “通用世界模型不是单一的生成器、模拟器、机器人动作模型或策略模型,也不是这些能力的割裂片段或简单线性串联,而是三种能力耦合在一起的闭环反馈系统。”朱军强调,行动不仅是模型的输出,还会改变环境、产生新信息,并进入下一轮理解、预测和…

2026世界机器人大会(WRC)于8月19日在北京亦创国际会展中心燃情启幕。相较往届,这场全球机器人领域规格最高、规模最大的行业盛会,已从“机器人秀场”进阶为人机共生、产需共融的“实战考场”。 浙江人形机器人创新中心有限公司(下称“浙江人形”)携NAVIAI核心机器人矩阵亮相智创馆C-105展位,集中展示同一技术基座下的产品体系与多场景泛化落地成果。展期吸引多批客户深度洽谈,获行业专家、媒体及观众高度关注。 多场景泛化落地的背后,是SPIRE技术体系对世界模型与具身模型的深度融合:世界模型构建环境认知、预测状态变化,具身模型动态适配、输出精准动作,两者协同让机器人完成从“听懂语义”到“做对动作”的跨越。正是基于这一技术底座,浙江人形在工业、商业、家庭、数采等多场景全面开花。 工业场景中,浙江人形率先跨越“单机演示”阶段,由3台NAVIAI轮臂机器人协同展现任务级全链路集群作业范式,统一拆解装配任务、分配工序,保障0.03毫米级操作精度,完整覆盖拆垛、分拣、搬运、装配全链路。 该方案是在汽车产线、海外工厂等多场景实践的能力复刻,印证了人形机器人从“单机负责单一功能”跃升为“集群重构柔性产…
Insider Brief Veeda AI has raised more than $90 million in seed funding to build world models designed to give robots virtual environments where they can learn through trial and error, The Logic has reported. The Canadian startup, founded by former Nvidia AI researchers Sanja Fidler, Zan Gojcic and Huan Ling, is backed by Radical […]
Current world models like Sora or Genie only simulate physics and ignore what people think, want, or feel. The new "Mental World Modeling" framework adds mental variables like beliefs and intentions. Even weaker language models using this approach outperform stronger models without mental modeling. The biggest bottleneck: predicting how physical and mental states change together. The article World models that ignore human beliefs predict the wrong actions, new research shows appeared first on Th…
当前,大模型主要通过预训练和后训练获得能力,上线后参数通常不再更新,只能借助上下文、知识库或重新训练适应变化。如何让 AI 在部署后持续学习,正成为行业需要突破的基础问题。 8 月 18 日,红杉资本旗下播客《Training Data》发布了一期围绕 AI 持续学习的对谈,受访者为强化学习奠基人、2024 年图灵奖得主 Rich Sutton。此外,参与访谈的嘉宾还有 Sutton 的前学生、Oak Lab 联合创始人 Khurram Javed,对谈由红杉资本合伙人 Sonya Huang 和 Alfred Lin 主持。 访谈从 Sutton 2019 年发表的《苦涩的教训》谈起,延伸至合成数据、世界模型和持续学习,并讨论智能体如何从真实经验中形成抽象概念。Sutton 认为,部署后便停止改变的模型,还称不上完整的智能系统。大模型虽是重要突破,却主要解决了语言能力,只覆盖智能的一部分。 基于这一判断,Sutton 正将持续学习推进为一条完整的技术路线。就在上个月,他与 Khurram Javed 创立 AI 初创公司 Oak Lab,并提出以经验学习、时间抽象和规划为核心的 Oa…

8月19日,以"人机共生、产需共融"为主题的2026世界机器人大会在北京正式开幕。作为全栈具身智能大模型公司,超维动力携全球首个人形机器人自主乒乓球完整对局成果、SMASH 2.0高动态人形乒乓系统、KAI世界模型、全球最高117个全身自由度的KAIBot高拟人本体、KAI Hand超高自由度灵巧手与KAI Halo第一视角数采头环等全栈矩阵登场,为观众带来一场"看得见、玩得到"的具身智能硬核体验。 SMASH 2.0:全球首个自主完整对局,现场开放人机对打 此前,超维动力已正式发布全球首个人形机器人自主乒乓球完整对局,标志着高速动态场景下"感知—决策—控制"全链路自主能力的里程碑式突破。本届大会,超维动力在展台现场搭建标准乒乓球交互体验区,全面展示基于自主研发的新一代SMASH 2.0系统——从视觉感知、轨迹预测、动作规划到全身控制的闭环串联能力。 乒乓球速度极快、旋转多变、落点难测,留给机器人的反应时间仅在毫秒级,真正难的不只是挥臂,而是肩、肘、腕、腰、腿多关节在极短时间内的协同配合。SMASH 2.0融合高速视觉感知、实时轨迹预测、全身运动控制与具身智能决策,在毫秒级窗口完成来球…
Florent Delgrange won the Best Blue Sky Paper Award at AAMAS 2026 for his work Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments. We caught up with him to find out more about his vision for agent learning. What is the topic of your Blue Sky Ideas paper and […]
当机器人开始理解世界,通用智能才进入物理空间。 作者丨吴思梦 陈嘉欣 编辑丨岑 峰 过去几年,机器人解决的问题是“如何更精准地执行动作”。 在工业生产线上,机器人可以完成焊接、搬运、装配等大量重复任务,但这些能力建立在一个前提之上:环境、流程和目标都已经被提前定义。机器人并不需要真正理解世界,只需要按照预设路径完成任务。 但当机器人开始走出工厂,进入物流、服务甚至家庭场景,问题正在发生变化。面对一个没有见过的物体、一个没有训练过的任务,机器人能不能理解环境、预测变化,并找到完成任务的方法? 这成为通用机器人走向下一阶段必须回答的问题。 大模型的发展为机器人提供了新的可能。语言模型证明,人工智能可以通过大规模数据学习复杂规律,并获得跨任务泛化能力。但机器人面对的并不是语言世界,而是真实的物理世界。 因此,具身智能领域正在探索一个更底层的问题:大模型究竟应该如何赋能机器人? 目前行业的主流方向是 VLA。它让机器人通过学习大量人类行为数据,将视觉、语言和动作统一起来。但 VLA 的核心仍然是模仿——模型学习人类过去如何行动。 问题在于,如果机器人面对一个从未见过的新任务,仅依靠已有动作经验…
This year's World Robot Conference staged a robot 'wedding' on Qixi and unveiled five industry trends: robots moving from technology validation to product validation, world models focused but not as hot as expected, tactile sensing becoming the new focus of dexterous hands, exoskeletons showing initial promise, and joints becoming component suppliers' new battleground.
Veeda AI, a startup led by a team of former Nvidia Corp. researcher and renowned computer scientist Sanja Fidler, has taken its bow on the main stage after raising $90 million in a seed funding round today. The round, which was first reported by The Logic, was co-led by Khosla Ventures and Radical Ventures, is […] The post Sanja Fidler’s world model startup Veeda AI raises $90M in seed funding appeared first on SiliconANGLE.
Robots need policies that can adapt to their sensors, environments, and tasks while running on onboard computing hardware. World models offer a foundation for... Robots need policies that can adapt to their sensors, environments, and tasks while running on onboard computing hardware. World models offer a foundation for learning physical interactions, but their size can make on-device deployment difficult. This changes with the new NVIDIA Cosmos 3 Edge. Cosmos 3 Edge is a 4B omni-model (with a 2B…
