arXiv:2608.26673v1 Announce Type: new Abstract: Large pretrained vision-language-action models dominate modern robot-manipulation benchmarks, but it remains unclear how much model scale is necessary for strong language-conditioned control, or whether fundamentally different control architectures can remain competitive at much smaller parameter budgets. We present PredVLA, a language-conditioned predictive-coding policy with only 0.68 million trainable network parameters and no robot-data pretrai…
arXiv:2608.26669v1 Announce Type: new Abstract: Testing of commercial Advanced Driver Assistance Systems is essential to ensure safety and compliance during type approval and in service operation. However, proving ground scenarios may not reflect real world driving complexity, while geo fencing can require manufacturer collaboration and limit assessment independence. This work presents a methodology for independently testing Assisted Lane Change systems on public roads. A campaign on the A31 Fre…
arXiv:2608.26645v1 Announce Type: new Abstract: Vision-Language-Action Models~(VLAs) have demonstrated significant promise in generalizing to complex, long-horizon robotic manipulation tasks. However, their performance remains brittle, as they are typically trained on trajectory-monotonic, failure-free demonstrations. This reliance on ``perfect" data leaves them unable to recover from common execution errors, such as a missed grasp, a dropped object, or an unexpected collision. In this paper, we…
arXiv:2608.26622v1 Announce Type: new Abstract: Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. Compliance enables safe and adaptive contact, yet the intrinsic viscoelasticity of soft polymers leads to stress relaxation and a continuous decay of grasping force during holding. Inspired by human grasping, which combines phase-dependent stiffness regulation with continuous sensing and feedback, this paper presents an integrated structure--perception--le…
继创新形态的阔折叠手机取得了市场的广泛好评后,华为延续阔屏设计,带来全球首款阔直板旗舰HUAWEIPuraXView。HUAWEIPuraXView薄至6.68mm,轻至201g,拥有7000mAh超 ...查看全文
SOLO 框架针对感知型人形机器人长时程行走的误差累积问题,用傅里叶编码查询重建保留地形边界细节,并以轨迹感知蒸馏传递未来误差惩罚;压力地形测试通过率 97.5%,实机仅凭深度相机与本体感知零样完成 1.5 公里户外路线。AI summary
arXiv:2608.26578v1 Announce Type: new Abstract: This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails (e.g., causing the robot to grasp with a specified positional o…
arXiv:2608.26545v1 Announce Type: new Abstract: Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors. A common approach to combat such catastrophic forgetting is to train on new task data with a replay buffer of previously learned task data. Although this buffer is commonly sampled randomly from all prior experiences, we show that a small set of these experiences contributes greatly in anchoring past performance. We c…
arXiv:2608.26505v1 Announce Type: new Abstract: The Poppy Humanoid is an open-source, low-cost robot suitable for research and education in artificial intelligence. However, we are unaware of any published methodology that achieves reliable, unassisted bipedal locomotion on the standard Poppy hardware. This paper contributes a functional closed-loop walking controller for Poppy, based on the linear-quadratic regulator (LQR) framework for trajectory tracking. Starting with data collected from ope…
arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important concern when agents must operate continuously in the real world. Most state-of-the-art methods are still developed in synchronous simulators, where the environment waits for the agent to act and inference ti…
arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionable object representations within the latency budget of a physical control loop. Neural 3D reconstruction methods have demonstrated high-quality view synthesis, but their real-time viability across the compute platforms on which laboratory robots actually run remains poorly characterized. In this work, we present a system…
arXiv:2608.26314v1 Announce Type: new Abstract: Steering-based planners require solutions to state-to-state boundary value problems, which can be inaccessible for nonlinear platforms. Forward propagation evades the steering requirement, but the finite-sample behavior of the associated planners remains uncharacterized and their implementations underperform in practice. This paper develops a propagation-based kinodynamic planner with deterministic finite-sample near-optimality guarantees. We work …
arXiv:2608.26273v1 Announce Type: new Abstract: Static shape estimation of co-manipulative continuum robots (CCRs) is challenging because the continuum arms and manipulated flexible object form a closed chain that must satisfy both static equilibrium and geometric loop-closure constraints. This paper presents a constraint-aware physics-informed neural network (PINN) for static shape estimation of a tendon-driven CCR modeled using the geometric variable strain formulation. The proposed method inc…
arXiv:2608.26239v1 Announce Type: new Abstract: Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visua…
2026 年,智能体将在企业级应用中取得哪些实质性突破?点击下载"《2026 年 AI 与数据发展预测》白皮书,获悉专家一手前瞻,抢先拥抱新的工作方式! 企业 AI 的第一代重点,是把 agents 构建出来;下一代的重点,则是把它们稳定地运营到规模化。 Cortex Agents 一直为在 Snowflake 内构建 AI applications 提供 managed runtime。过去一年里,我们看到客户从部署第一个 agent,逐步走到开始管理企业级部署。这种转变暴露出一组全新的运营挑战,主要集中在 orchestration、execution 和 governance 上。 现在,我们正在为 Cortex Agents 扩展一批新能力,目的是降低企业在大规模构建、部署和运营 AI agents 时面临的操作复杂度。 快速总览 Build Coding Agent(即将进入 public preview):基于与 CoCo 相同的 runtime 构建一个托管式 coding agent,并把它部署到任意应用中;Skills Package(即将进入 public prev…

发表在《科学》期刊上的一项研究分析了加拉帕戈斯群岛(Galápagos)珊瑚长达千年的记录,发现随着地球变暖,厄尔尼诺现象正在加剧,东太平洋的厄尔尼诺-南方涛动(ENSO)变率比工业化前时代高出了约 36.5%。这些发现表明,气候变化已经在放大全球最重要的极端天气来源之一,并可能对生态系统、基础设施和人类社会带来日益增长的风险。ENSO 是导致年际气候极端事件的主要原因,它与干旱、洪水、野火、珊瑚白化事件以及对农业和人类健康的影响相关。该现象源于热带太平洋与大气之间的复杂相互作用。近几十年来发生了数次异常强烈的厄尔尼诺事件,其变暖范围覆盖了热带太平洋的大部分区域。

8月28日,阿里巴巴集团旗下高德正式发布首个无长程依赖的万帧级流式3D重建模型ABot-Recon。
作者 | Jackie Calapristi 编译 | 蔡芳芳 编者按: 当 AI 编程工具从“帮工程师写代码”进一步走向能够自主规划、实现、测试和排查问题的 Agent,软件工程真正发生的变化,可能已经不只是“写代码更快了”。Calendly 最近公开了其内部正在运行的一套 Agentic Engineering 实践:Agent 会先审查 Jira 需求、拆解任务,再并行写代码、跑测试、做 QA,最终由人类完成 Review 和 Merge;在 IAM 团队,仅一周半时间,Agent 就生成了 64 个 PR,完成了一个大型解耦项目约 30% 的工作。更值得关注的是,随着执行能力被进一步自动化,Calendly 发现工程研发的瓶颈再次发生了转移——从写代码,到需求对齐,再到如今的“判断力”。Calendly 是一家成立于 2013 年的日程安排 SaaS 公司,核心产品通过连接用户日历,自动处理会议时间协调、预约和工作流等问题,被大量企业和个人用于销售、招聘、客户服务等场景。正因为它本身是一家拥有成熟产品、工程体系和生产环境的 SaaS 公司,Calendly 团队这次分享的 Ag…

过去一个周末,全球AI开发者都在追问同一个问题:突然出现在 OpenRouter 上的神秘模型Ox Alpha,究竟来自哪家实验室? 它没有公布开发者,没有披露参数规模,也没有给出技术报告,只以“匿名模型”的身份开放使用。 但在短短几天内,Ox Alpha迅速冲上OpenRouter热门榜单,并凭借代码生成、复杂推理和长时间Agent任务中的表现,引发了大量猜测。 “牛来”模型谜底揭晓,是 GLM-5.3-Flash 现在,谜底终于揭晓。 昨天晚上,智谱正式确认:Ox Alpha 正是其最新发布的 GLM-5.3-Flash。过去一周在海外开发者社区引起轰动的“神秘模型”,真的来自智谱。 Ox Alpha最早于8月20日匿名登陆OpenRouter。 在身份揭晓前,OpenRouter对它的描述是:一款面向代码、持续智能体工作和生产负载设计的推理模型,适合处理长期软件工程、复杂推理,以及结合文本和视觉信息的工作流。 它可以接收文本、图片和视频输入,输出文本,匿名测试版本提供104万Token上下文。 OpenCode 还曾宣布向用户免费开放一周,并称模型提供方准备了每天100万亿Tok…

2026 年 8 月 28 日,全球 XR 眼镜领导品牌 VITURE 携手《影之刃零》开发商灵游坊(S-GAME),于科隆游戏展期间举办的腾讯十周年全球艺术展上,正式揭晓收藏级联名之作——VITURE ×《影之刃零》Phantom Beast「幻影猛兽」XR 眼镜。 这不仅是一款联名产品,更是一件可以佩戴的典藏作品:它将《影之刃零》独树一帜的「功夫朋克」美学融入 VITURE 旗舰级 XR 技术,让玩家透过 174 英寸沉浸巨幕,真正踏入刀光剑影交织的武侠世界。 联名眼镜即日起于 VITURE 京东自营旗舰店独家开启预购,首发价 4,299 元,并将于 10 月 29 日与《影之刃零》同步发售、正式发货。预售期间下单,还可获赠独家珍藏版游戏地图卷轴。 “武侠是伴随我成长的重要文化元素,而《影之刃零》对武侠世界的全新诠释,正是我期待多年的样子。”VITURE 创始人姜公略表示,“如果说《影之刃零》创造了一个前所未见的功夫朋克江湖,那么 VITURE 想做的,就是成为玩家进入这片影境的入口——不只是看见游戏,而是真正置身其中。” 《影之刃零》制作人梁其伟表示:“VITURE ×《影之刃零…

Chinese surgeons last month completed the world’s first commercial surgery to implant an invasive brain-computer interface (BCI) device in a patient with a spinal cord injury, a milestone in the neurotechnology development race with US companies like Elon Musk’s Neuralink. This month a Chinese company, state-owned PICC Property and Casualty, launched the world’s first commercial insurance policy covering BCI implantation surgery. September will see the first undergraduates in China to major in..…

A federal judge has called the Department of Defense’s designation of Anthropic as a national security supply-chain risk “illegal and baseless.”
英伟达收购Hugging Face引发了开发者群体担忧,对于中国开源生态来说,风险亦在增加。

站在 2026 年 8 月北京世界机器人大会(WRC)的展馆中心,如果你闭上眼,听到的不再是两年前那种对“机器人像人一样跳舞”的惊呼,取而代之的,是从业者们在心里拨动算盘,计算 ROI(投资回报率)的声音。 具身智能正在经历一场从“秀肌肉”到“翻账本”的集体变轨。 回看这半年的轨迹,这种趋势演进得极具层次感: 5 月的维也纳 ICRA,学术界和产业界第一次如此默契地把所有火力都集中在了 VLA(视觉-语言-动作)模型上,那是具身智能的“物理学补课期”; 7 月初的上海 WAIC,200 多家具身智能企业同台竞技,集体把演示从“后空翻”换成了“拧螺丝、叠衣服”,那是行业的“场景落地焦虑期”。 7 月底的悉尼 RSS,开源硬件、仿真工具链和遥操作数据方案成为主角,行业开始为“大脑聪明、身体廉价、数据贫瘠”搭建公共底座,那是产业的“基础设施奠基期”。 而在 8 月的北京 WRC,所有话题似乎都围绕工程约束和商业底线展开。在这里,没人关心你的模型是否有哲学意义上的“意识”,大家只盯着三个指标:多少钱?多稳?跑一个月能省几个人的工资? “以前我们聊的是什么时候 AGI,现在我们在聊这台机器人跑一…

谷歌发布号称“迄今最强”语音转文本模型 8月26日,谷歌发布新一代语音转文本模型 Gemini 3.5 Transcribe,可以在转录过程中自动处理“嗯”“啊”等语气词、口误和重复表达,并补充标点、大小写及文本格式。 谷歌将其称为“迄今最精确的语音转文本模型”,谷歌 CEO Pichai 在x上宣布了该模型。 与只追求逐字记录的传统语音识别模型不同,Gemini 3.5 Transcribe 希望直接将原始语音转换为更接近成稿的文本,减少用户后续整理会议记录、采访速记和通话内容的工作量。 例如,当用户说“我们周二开会——不,改成周三”时,模型能够理解后半句是对前面内容的修正,并在最终文本中保留正确结果,而不是机械记录整段口误。 Gemini 3.5 Transcribe支持85种以上语言和地区变体,并能够自动判断当前使用的语言。用户不需要提前设置语种,即使在一句话或者同一段对话中切换语言,模型也可以继续转录。 这项能力主要面向跨国会议、多语言访谈、客服通话以及同时夹杂中英文专业术语的场景。 在真实录音中,影响转录准确率的往往不只是口音,还包括背景噪声、多人同时发言和专业词汇。谷歌称,…

Cloudflare" 最近详细介绍了" 其如何利用 AI 将内部工程标准从被动文档转变为在软件开发生命周期中主动执行的控制系统。报告显示,自 2026 年初以来,该公司的 AI 代码审查工具已识别出近 23 万处不符合工程规范的问题,其中近 1.6 万处导致审批被驳回。Cloudflare 还将同样的方法应用于技术设计和事故报告,使用一个叫作 Cloudflare Codex 的中央知识库作为其工程标准的单一事实来源。 这一转变并非简单地把 AI 用于代码审查,而是将工程知识改造为机器可读取、可强制校验的形态。通过结构化的 RFC 文档定义标准,约束被分为建议(SHOULD)与强制(MUST)两个等级,并赋予明确的所有者和生命周期状态。新规范上线时可先仅提供建议,后续再过渡到能够阻断代码变更的强制管控规则。由此形成“指引建议 → 观察 → 强制执行”的递进模式,使治理成为开发工作流的一部分,而非工程师需要单独查阅的内容。 Cloudflare 将该模型应用于开发的多个阶段。AI 可以在开始实现之前审查技术规范,在开发期间依据同一标准检查代码,并在事后评估事故报告。这形成了一个潜在的强…

游戏行业增长逻辑正在从买量转向玩家社区与内容生态。亚马逊广告通过Twitch、Prime Video等布局,试图打通内容场景与受众信号,重塑开放互联网的广告投放逻辑。本文深入解析亚马逊广告的扩张路线,揭示其如何连接内容与消费者,为品牌提供新的增长路径。 一 今年ChinaJoy上,游戏公司谈增长的方式变了。 前几年,展馆里最常出现的词还是买量、ROI。现在,被反复提及的是全球化、PC与主机、创作者以及玩家社区。 手游市场正在进入存量阶段。Sensor Tower发布的《2026年游戏市场报告》显示,2025年全球手游下载量约为520亿次,继续下滑;内购收入约820亿美元,同比增长只有1%左右。 另外,2025年全球游戏产业规模已接近2000亿美元,其中PC与主机游戏贡献约950亿美元,占据半壁江山。中国发行商贡献了全球约三分之一的游戏数量,但仅占全球收入的4.6%,形成了数量与收入之间的巨大落差。 这对游戏厂商的增长提出了一个更难的问题。 以前的增长,可以被拆成一套数学模型:买进用户,计算转化,控制成本,在生命周期内收回投放。进入PC、主机和全球发行市场后,这套模型会显得有些力有不逮。…
