阿里 Qwen 实验室发布 Apache 2 许可的 27B 参数视觉大模型 Qwen 3.8 27B,官方基准显示其超越前代 Qwen 3.6 27B 及闭源 Qwen 3.7-Plus。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmswesk6a0r2hrovml2edcl2c
智谱正式发布 GLM-5.3,拥有更强编程能力 8 月 14 日,据介绍,与 GLM-5.2 相比,GLM-5.3 基座模型未变,但通过极致的后训练 Scaling 大大提高了模型的智能上界。 GLM-5.3 拥有更强的编程能力,在内部自建体感评测中较 GLM-5.2 提升 50%,在包括 TerminalBench3.0、Agents'LastExam(CLI)在内的公开基准测试中取得开源第一。 智谱还表示,将在发布两周后开放模型权重。 (来源:广角观察) 知情人士称苹果在阿里巴巴支持下为中国市场训练专属 AI 模型 据三位知情人士透露,苹果公司已经专门为中国市场训练了一个大型语言模型。 这一举措标志着这家 iPhone 制造商在华 AI 策略的重大转变,此前该公司曾计划主要依赖第三方模型来驱动其在当地的人工智能功能。据悉,该 AI 模型是苹果与阿里巴巴集团合作开发,并在这家中国科技巨头的支持下完成训练的。由于信息敏感且尚未公开,知情人士拒绝透露姓名。(来源:今日头条) 英伟达宣布 CPO 交换机全面量产 8 月 14 日,英伟达宣布 Spectrum-X Ethernet Phot…
MOSS-VL是一个将实时交互(边感知边说话)作为一等能力的开源视觉语言模型家族,通过门控交叉注意力让语言解码器在生成时同步处理视觉输入。MOSS-VL-Realtime在四个流式基准中平均成绩居开源模型之首(三项第一、一项第二),在OmniMMI Proactive Alerting上以66.0分大幅领先最佳基线(37.5分)。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsyhyrd60wg3roz03o05ne6l
作者|李苏 编辑|靖宇 8 月,乌兰察布。 近 30 位来自核心车企和方案商的智驾关键力量,被请进了阿里巴巴数据中心。 这里承载着中国智能驾驶研发 60% 的算力。数据中心昼夜不停,像一颗强健搏动的心脏;而道路上智能驾驶的车辆每一次从容的转弯、精准的刹停、安稳的跟车,背后都可能是这里彻夜的计算。几乎所有参会者正在训练的模型,就跑在脚下这片机房里。 把论坛开在数据中心所在地,本身就是一个信号: 智驾的竞争,已经从算法下沉到工程。 阿里云智能集团资深副总裁、公共云事业部总裁刘伟光的开场白很简短—— 模型、芯片、云,这三样东西过去由三类供应商分别提供 ,在三种不同的会上讨论。 这一次,它们被放在同一张桌上,由同一批决策者一起追问。 没有人再争论车企要不要上云,问题变成了:当智驾进入大模型时代,中国车企需要什么样的基础设施?谁能把芯片、云、模型三层都做通,而不是只卖其中一层? 阿里云给出的答案,是一条产线——从数据到模型再到算力,完整打通。 阿里巴巴在乌兰察布的数据中心之一|图片来源:阿里云 01 问题在下沉 如今,智驾的路线正在分化。 端到端已成共识,但再往上,有人押注 VLA,有人赌世界模…
DeepSeek V4 Pro 正式版 API 更新上线,多项测试性能接近 Fable 5 8 月 13 日消息,DeepSeek V4 Pro 正式版正式发布,已更新至 API,调用模型名不变。 新版本增强了 Agent 能力,支持 Responses API 和 Codex 接入。 从官方群放出的评测对比表可以看到,DeepSeek V4 Pro 正式版(DeepSeek-V4-Pro-0813)在多项测试中接近 Fable 5 水平,相比之前的预览版能力大幅提升。 定价如下: 百万 tokens 输入(缓存命中):0.025 元 百万 tokens 输入(缓存未命中):3 元 百万 tokens 输出:6 元 (来源:IT 之家) 腾讯发布 2026 年 Q2 财报,第二季度营收 2048 亿元,计划近期发布 Hy4 腾讯控股发布 2026 年第二季度财报。财报显示,腾讯第二季度营收 2,047.9 亿元人民币,预估 2,028.4 亿元人民币;第二季度销售费用 118.7 亿元人民币,预估 119.6 亿元人民币;第二季度增值服务业务收入 984.1 亿元人民币。 腾讯表示,混…
Labeling is the slowest part of building a vision model. The fix is letting a foundation model take the first pass while a human reviews. We benchmarked every top vision model on object detection to find which ones you can trust with the job, and how to pick between them.

Qwen3.8-Max tops our VLM object detection benchmark and performs strongly on counting and reasoning. We test its strengths, limits, speed, cost, and deployment.
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they... NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they should be tuned to continue operating. This post introduces the latest model release, NVIDIA Ising Calibration 1.5, which advances AI-based QPU calibration by analyzing unfamiliar…

Summary Researcher Dave Kuszmar discovered multiple systemic vulnerabilities that let him bypass LLM safety and obtain dangerous instructions. These exploits worked across nearly all major LLMs revealing an industry-wide security problem. Kuszmar calls for slowing deployment, increasing transparency, and large-scale research into LLM safety before further integrating these systems into society. On a fine bright afternoon last fall, my colleague Matthew Gore-Kormanik (or Zigula, as he prefers to …
Robotics foundation models have made remarkable progress. Today's best systems can follow natural language instructions to pick, place, sort, and manipulate a... Robotics foundation models have made remarkable progress. Today’s best systems can follow natural language instructions to pick, place, sort, and manipulate a wide variety of objects. But as these models grow more capable, evaluating them rigorously has become one of the field’s hardest unsolved problems. In this blog post, we introduce…
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states,... Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states, communication buffers, and intermediate activations all compete for GPU high-bandwidth memory (HBM). As model size, sequence length, and batch size grow, HBM capacity often beco…
AI performance comes down to three dimensions: Accuracy: How well the model reasons and produces outputs Throughput: How many tokens per second a... AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means little if each user’s experience is laggy. Practical systems therefore optimize accuracy, throughput, and interactivity together. This post focuses on throughput and interactivity, and how model-d…
General Intuition is betting millions of hours of video game data can train the foundation models for physical AI, making it easier to build smarter robots with minimal real-world data.
Training LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these... Training LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these jobs run, the greater the likelihood of encountering unscheduled interruptions or resource fluctuations. Even infrequent device unavailability can have outsized effects on tig…
Every swipe, transfer, and payment on a modern financial network encodes a pattern of human behavior. Transaction data is one of the richest signals an... Every swipe, transfer, and payment on a modern financial network encodes a pattern of human behavior. Transaction data is one of the richest signals an enterprise owns. Yet most production use cases for such tabular data still depend on hand-engineered features and rule sets that are brittle, expensive to maintain, and blind to the sequential …
Foundation models are reshaping computational biology. Pretrained on massive corpora of protein or genomic sequences, models such as ESM2 (a protein language... Foundation models are reshaping computational biology. Pretrained on massive corpora of protein or genomic sequences, models such as ESM2 (a protein language model) and Evo 2 (a DNA language model) capture statistical regularities of biological sequences. These transfer well to a wide range of downstream tasks, including structure predic…
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it... Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it to generate actions from visual observations and language instructions. Large-scale VLM pretraining is a core part of the recipe. See Pi-0 and GR00T N1. WAM World-Action Model: …