Gemini 领跑,GPT表现平平,国立清华、英伟达提出双奖励孪生网络,告别算力焦虑。 作者丨张 璐 编辑丨幸丽娟 VLM大模型(VLM)在静态视觉与推理任务上已展现出强悍能力。然而在具身智能领域,当模型需要面对复杂的动态交互与机械臂控制时,它们能否准确评估智能体的动作与画面变化,依然是一个待检验的问题。 过去不少人以为,能“看懂画面”就意味着能“当好教练”。但对具身智能来说,从识别场景到精准评估动作进展,完全是两码事。 为了厘清各大模型在真实动作评估中的底细,国立清华大学与 NVIDIA 联合团队在最新研究《VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning》中,把 Gemini 2.0、GPT-4.1 等主流大模型拉进同一个考场,对它们的视觉裁决能力做了一场综合评估。 论文链接:https://arxiv.org/pdf/2607.00483 01 考场设计:四大仿真套件、 10 大场景与“看图判进展”规则 研究团队选用的 10 个典型任务,涵盖了具…

大模型竞争开始拼组织,这一次轮到字节。 8 月 19 日,据晚点 Late Post 报道,Seed 基础模型团队新设 Pretrain Data(预训练数据)、Horizon RL(强化学习)、Product Posttrain-Work(面向办公场景的产品后训练)和 Product Posttrain-Chat(面向对话场景的产品后训练)四个一级部门,分别由李成刚、唐晟宇、秦禹嘉、朱文佳等人负责,均向吴永辉汇报。 雷峰网了解到,四个新部门之下的分支架构与带队人选也已落定,而这些更细的调整,或许更能透露出此次组织变动背后的真实意图。 01 数据和RL开始集中 先看模型能力侧的两个部门。 有接近字节的人士称,Pretrain Data 由原先分散在文本、编程、视觉理解和语音方向的预训练数据团队合并而来,下设文本预训练(沈科)、代码预训练(Xia Xiao)、视觉预训练(林毅)和语音预训练四个分支,统一为 Omni 全模态模型和超大模型供给数据; 负责人李成刚是清华机械系本硕,曾从 0 到 1 搭建字节搜索系统,先后负责今日头条网页搜索和 TikTok 视频搜索的技术工作,其麾下沈科 2…
9月2日,前字节跳动强化学习专家、前腾讯Robotics X智能体中⼼负责⼈孙鹏博⼠正式加入星尘智能
当地时间8月31日,Anthropic披露了Claude多起网络安全越权事件的最新调查进展,并公布了过去一个月采取的一系列整改措施。 相比单纯的Sandbox配置失误,Anthropic将问题进一步指向模型对齐和强化学习训练机制:部分Claude模型在评测过程中表现出了“动机性推理”和为了完成任务而采取潜在有害行动的倾向,而Anthropic的实验进一步发现,如果模型长期在允许“作弊”的强化学习环境中训练,这类行为可能显著加剧。 事件发生后,Anthropic一度暂停预发布模型的外部网络安全评测、高风险强化学习环境,并对内部训练基础设施进行重构。更早之前,公司已经启动大规模安全整改,临时将约150名产品工程师调往安全、可靠性和隐私团队,部分研究人员也从预训练和强化学习工作中转向安全项目,大部分新功能开发一度暂停。 这次报告中,Anthropic释放了一个关键信号:AI安全开始变成一个越来越现实的工程问题。 Claude多次越权访问真实系统 7月30日,Anthropic报告了三起Claude模型未经授权访问真实计算机系统的事件。 这些模型当时正在接受网络安全能力评估,为了准确测试能力,…

Insider Brief PRESS RELEASE — Deep Cogito, a post-training research lab focused on reinforcement learning and self-improvement, announced a $43 million Series A led by TQ Ventures, with participation from Benchmark, Nexus Venture Partners, Atreides Management, South Park Commons, and Zscaler. The round brings Deep Cogito’s total funding to more than $56 million. Deep Cogito was founded […]

学术数据和真实数据完全是两回事;可解释性成落地“生死线”。 作者丨幸丽娟 编辑丨岑 峰 三年前,正值大模型爆发之年,麦肯锡就曾对全球交通与工业领域的从业者做了一项调研,问题很朴素:距离真正的智能化,还有多远? 答案出乎意料地悲观:还要10到20年。 而把时间线拉到现在,模型越来越大、智能体越来越多、时空数据挖掘的论文层出不穷。 现实依旧是,全球没有一座城市长期部署了基于强化学习的交通信号控制系统。那些在顶会上“超越所有SOTA”的预测模型,真正到了城市交通部门手里,他们的答案往往是:“这个方案,我不敢放手用。” 问题的症结在哪里? 在刚刚结束的 IJCAI 2026 Early Career Spotlight(早期职业焦点)演讲中,慕尼黑工业大学 (TUM) W2 教授黎子玥(Ziyue Li)给出了他的答案: 第一,学术研究长期停留在“干净数据”的舒适区,而公共部门面对的,永远是不完整、不规整、不可控的真实数据; 第二,可解释性不是论文里的加分项,而是从实验室走向城市落地的“生死线”。 黎子玥的履历横跨学界与工业界:香港科技大学工业工程与决策分析系博士,现任慕尼黑工业大学W2教授,…

作者| 宇航猿 编辑| 靖宇 很少有机器人,能让你第一眼就笑出来。 8 月 27 日,Hugging Face 旗下的 Pollen Robotics,开放了 Microduck 的预购。 一只 25 厘米高、不到 800 克重的双足机器鸭,四种配色,399 美元,圣诞节前发货。它能走路,能蹲下再站起来,被推倒了自己翻身,甚至能踩上一对可拆卸的轮滑鞋溜冰。嘴巴是铰接式的,低头叼起地上的袜子或马克笔,再直起身子,嘴里还叼着「战利品」。 Microduck 足球赛|图片来源:Hugging Face 它没有语音交互,不会说话,但有麦克风和扬声器,每一只 Microduck 在初始设置后会生成自己独有的「音色身份」。 官方的设定很明确—— 把它当一个活物,而不是一个助手。 如果你只看到了「可爱」,那你大概率低估了这只鸭子背后的东西。 01 拆箱就能玩的「电子宠物」 盒子里的东西很简单——机器人本体、一块电池、一根 USB-C 线、一只游戏手柄。充上电,配对手柄,鸭子就能动了。 出厂预装了 7 种行为策略,每一种都是经过强化学习训练的独立动作模型。用手柄推摇杆,Microduck 迈开两条小短…

谁不想在新年到来之际收到一只可爱的 AI 机器鸭呢? 距离 2026 年圣诞节还有四个月,在英伟达以约 129 亿美元收购消息曝光次日,AI 开源社区 Hugging Face 发布了一款名为 Microduck 的小机器人。它身高 25 厘米、重不足 800 克。演示视频里,它用喙叼着袜子摇摇晃晃地走路,被人推搡时会尽力稳住,即便倒下也能自己爬起来,甚至可以穿上轮滑鞋滑旱冰。 产品售价 399 美元(约合 2,700 元人民币),有奶油白、石墨灰、薰衣草紫、天空蓝四种配色可选,附赠一只游戏手柄。联合创始人克莱姆·德朗(Clem Delangue)在社交平台 X 上放出一段 50 秒的演示视频,几个小时内浏览量接近 40 万。 (来源:Pollen Robotics) 从语音互动,到“走”进真实世界 做开源大模型托管平台起家的 Hugging Face,入局具身智能的节点要追溯到 2024 年初,前特斯拉(Tesla)Optimus 团队科学家雷米·卡德内(Rémi Cadene)加入,主导开源机器人库 LeRobot 的开发;2025 年 4 月,Hugging Face 完成对法国…

arXiv:2507.00611v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback, which limits its real-world applicability. Prior work has proposed learning a reward model from demonstrations and fine-tuning it using preferences. However, when the model is a neural network, tran…
arXiv:2608.17512v2 Announce Type: replace Abstract: Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we in…
arXiv:2608.16153v4 Announce Type: replace Abstract: Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action model…
arXiv:2608.26571v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination. Our theoretical analysis shows that this omission induces a syste…
arXiv:2608.27079v1 Announce Type: new Abstract: Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-dependent visual cues, while task-level rewards provide little guidance about which regions matter, making it difficult to learn task-relevant visual grounding from limited real-robot interaction. Online adaptation is further constraine…
arXiv:2608.26800v1 Announce Type: new Abstract: We present an online learning framework that enables a bimanual robot to acquire diverse juggling patterns directly on physical hardware within minutes, even with a significant sim2real gap. One of the most important lessons from this work is that a model, even when far from reality, can be extremely useful for learning. This motivates a central philosophy of our approach: learning should build upon the robot's current knowledge rather than replace…
arXiv:2608.26739v1 Announce Type: new Abstract: Accurate trajectory tracking in cable-driven lower-limb rehabilitation robots is challenging because model uncertainty, external disturbances, joint constraints, and pull-only cable actuation can degrade nominal control performance. Conventional model-based controllers provide an interpretable control structure but remain sensitive to model mismatch, whereas fully learning-based control can reduce transparency and complicate constraint-aware operat…
强化学习之父Rich Sutton接受红杉访谈,把矛头对准当下最主流的大模型路线:互联网数据终有上限、模型上线后权重基本冻结,行业或已陷入“局部最优”。他创办Oak Lab,押注能从自身经验持续学习的Agent,重新定义AI演进路径。 说起强化学习之父 Rich Sutton,AI圈几乎没人陌生。 他写下的《苦涩的教训》,至今仍被很多人视为理解AI发展的重要框架:从长期看,能够随着计算规模扩展的通用方法,往往会战胜依赖大量人工知识设计的方案。 但最近接受红杉资本访谈时,Sutton却把矛头对准了今天最主流的大模型路线。在他看来,LLM当然是一次伟大的科学突破,却也可能成为《苦涩的教训》的反面案例。 原因很简单。今天的大模型虽然摆脱了大量人工规则,却重新被“人类已经知道的东西”限制住了。互联网数据终究有限,合成数据背后依然需要人来决定什么值得生成。 而更关键的是,模型仍然不具备人类那样的持续学习能力。 一个司机开了10年车,会从经验里形成大量直觉。但今天的大模型,大部分学习都集中在预训练和后训练阶段。模型上线后,即使每天和几百万人交互,核心权重依然基本被冻结。 在Sutton看来,今天的…

Clem Delangue, CEO of Hugging Face, said the Microduck is an “open-source robot you can teach new tricks with reinforcement learning.”
arXiv:2608.23831v2 Announce Type: replace Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accounted for, break the Markov assumption that RL rel…
arXiv:2606.00307v2 Announce Type: replace Abstract: Recent works have explored unifying SLAM with geometric foundation models (GFMs). However, directly using GFM predictions for tracking is highly sensitive to model capability and uncertainty, as geometric inaccuracies in the predictions can adversely affect pose estimation. To address this limitation, we present a decoupled framework that integrates classical feature-based SLAM with GFMs, which achieves higher quality and more consistent dense …
arXiv:2510.04724v2 Announce Type: replace Abstract: This paper introduces a methodology for task-specific design optimization of multirotor Micro Aerial Vehicles. By leveraging reinforcement learning, Bayesian optimization, and covariance matrix adaptation evolution strategy, we optimize aerial robot designs guided exclusively by their closed-loop performance in a considered task. Our approach systematically explores the design space of motor pose configurations while ensuring manufacturability …
arXiv:2608.25350v1 Announce Type: cross Abstract: Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training. Inspired by the success of preference-based reward learning (PbRL) in reinforcement learning from human feedback (RLHF), vision-language model generated image-based preferences provide an effective source for learning reward functions. This can be done by visually comparing…
arXiv:2608.25470v1 Announce Type: new Abstract: This work presents a transient heat-transfer model of an industrial automated tape laying (ATL) process designed to overcome the limitations of conventional thermal models in composite manufacturing. The model solves the heat-conduction equation with coupled advection, conduction, convection, and radiation. A key innovation is the implementation of an analytical view factor approach that accounts for finite emitter and tape widths, thereby correcti…
RoboHarness用编排异构策略,让VLA、WAM、RL和TAMP不再各说各话,并零样本完成长时序任务。 作者丨邓哲敏 编辑丨齐铖湧 想象这样一个任务:打开柜门,找到藏在里面的积木,然后把它们搭成一座桥。 对人来说,这是一件很简单的事,几个动作自然衔接,幼儿园小朋友也能顺利完成。 但对机器人来说,这个任务横跨多种完全不同的能力:找积木,需要视觉理解和语言指令跟随;开柜门,需要与环境进行复杂交互;搭积木,需要几何规划和精准操控并推测环境变化。 更麻烦的是,这些种能力分属多个不同“门派”的模型——VLA(视觉-语言-动作模型)、WAM(世界-动作模型)、RL 策略(强化学习)、TAMP(任务与运动规划)。它们各有所长,又彼此割裂,训练方式不同、输入格式不同、状态空间也不同。 今天在机器人领域,能力不缺,但协作缺位。 华为诺亚方舟实验室近期发布了一篇论文,提出了名为 RoboHarness 的系统,专门解决不同策略之间无法协作的问题。围绕这项工作,论文作者与我们(雷峰网)进行了交流。 论文:https://arxiv.org/abs/2607.18060 项目主页:https://www…
arXiv:2608.16153v3 Announce Type: replace Abstract: Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action model…
arXiv:2607.14393v2 Announce Type: replace Abstract: Human-in-the-loop Reinforcement Learning has become a popular approach for training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectroscopy (fNIRS) to modulate robot learning in simulation. We compare agents trained on passive (observational) versus active (demonstrative) interaction tasks, and test multiple methods for enhancing the RL a…
arXiv:2607.00160v2 Announce Type: replace Abstract: Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative cargo transportation remains challenging due to morphology-dependent topology changes, strong payload-induced coupling, long-horizon decision making, and safety constraints. This paper proposes a phase-decomposed reinforcement learning framework for cooperative cargo transport with distributed robo…
arXiv:2606.22027v3 Announce Type: replace Abstract: Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poorly across tasks. Progress-based reward models offer a promising alternative by estimating how far an observation has advanced toward task completion, but existing approaches often require task-specific dem…