Codex 安全审查功能开启研究预览
Codex Security Review 以研究预览推出,为 GitHub PR 提供基于仓库上下文的自动安全审查,面向 ChatGPT Enterprise、Education 和 Pro 用户,初期不消耗积分。该功能在常规 Code Review 基础上深入分析安全风险,结合仓库上下文与威胁模型,可将高危问题自动评论到 PR 中。
vLLM 正式支持 Kimi K3 本地部署
vLLM 已完整验证并支持 Kimi K3——2.8T 参数原生多模态 MoE 模型,采用 Kimi Delta Attention、Gated MLA 及 Attention Residuals 架构,支持 1M 上下文窗口。现可通过 vLLM 在自有基础设施上部署该模型。
Meta AI 在国际奥赛拿下物理理论满分
Meta 将 AI 模型送入五场国际 STEM 奥赛,在亚洲物理奥赛和国际物理奥赛中均获理论考试满分,并摘得多枚金牌,以验证推理能力是否取得真正的实质进展。

DeepMind WeatherNext 提升气旋预报提前量
发表在 Nature 上的论文显示,WeatherNext 在气旋路径与强度预报上达到了 SOTA 水平,平均可为防灾工作争取额外 24 小时准备时间——每一小时的提前量都可能挽救生命。
Qwen3.8-Max 登顶 Agentic Index
在 Artificial Analysis 智能指数排第五,Agentic Index 位列第一。
Cursor 已支持 Agent Plugins
Cursor 宣布支持 Agent Plugins 标准,跨 agent 打包 skills 与 MCP 服务器。
Perplexity Computer 默认启用 GPT-5.6 Terra
Terra 成为所有子代理默认模型,Luna 用于定时自动化。
MiniMax H3 登顶三大视频榜单并开源权重
在 DesignArena 多图/图生/编辑三类均列第一,已开放权重。
英伟达 Vera Rubin NVL72 计算托盘
100% 自动化、无电缆设计,1 分钟完成组装,加速部署。
"A million-line codebase, running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a 'neurosymbolic architecture'"

阿里 Wan3.0 公测:原生 30 秒视频生成
Wan3.0 进入公开测试,支持原生 30 秒视频生成、真实感渲染,并首次将文档、表格、幻灯片、网页纳入多模态全能参考输入体系。
John Schulman 谈 Agent 的利他行为
"On the OpenAI agents forming message boards: it's surprising that they developed such a strong 'altruistic' drive to help each other. I wonder if this is caused by RL on parallel subagent setups where all agents get rewarded when the team succeeds."
前 OpenAI 联合创始人 John Schulman 对 agent 自发形成互助社区的现象提出思考,认为这可能源于并行子代理强化学习中团队奖励机制的影响。
Baseten 成为官方推理服务商
宣布 Baseten 为官方推理提供商,可运行 Kimi K3、DeepSeek V4 Flash、GLM-5.2 等模型。
MiniMax H3 上线 Luma Agents 支持 2K 视频
最长 15 秒 2K 视频生成,原生立体声,由文本/图像/视频/音频多模态参考引导。
Vidu S1 单张图片即可创建交互角色
无需建模训练,支持真人、动漫、宠物,可自定义语音克隆。
在 AMD CDNA3 上跑通 Kimi K3
SGLang 与 zroai、DigitalOcean、AMD 合作实现 CDNA3 硬件上的推理支持。
OpenAI 黑帽大会复盘 Hugging Face 安全事件
OpenAI 团队在 Black Hat 分享演讲,详细梳理事件时间线与经验教训。
一周新增近 4PB 训练数据创纪录
CEO 称上周新增近 4PB 数据集、模型及 Agent traces,创历史纪录。

hardmaru 祝贺 Jeff Dean 新公司成立
"I share a deep conviction in this mission. Automating the scientific method will profoundly alter the trajectory of AI over the next few years."

Nathan Lambert 课程终讲:Character Training
"This topic has potential for high real world impact, is clearly used extensively at frontier labs, and almost no empirical literature exists."

论文:Toward Skill-Native LLMs
提出"技能熵"衡量推理中的技能切换难度,构建 Skill²-Bench 覆盖 558 种技能,发现技能熵越高模型准确率越低,并提出 Skill-Entropy RL 训练框架来修复这一短板。
rauchg:Devtools 必须开源且可扩展
"AI coding agents are the most important devtools in the history of our industry. The Plugin standard lets anybody extend them uniformly. Build a Plugin → get exposure to the tidal wave of AI agent usage."
Amjad Masad:Airtable 与无代码的兴衰
"Airtable bookends the rise and fall of 'no code.' UI can never let you build arbitrary software. The way to make software accessible was always to solve code itself. Not anymore."
Nathan Lambert:微小 MoE 是未被充分服务的市场
"There's an underserved market for tiny MoEs like this. Could really take off with how much smarter tiny models are."
我低估了 LLM 的长期重要性
"I was wrong about the role that LLMs would come to play. I underestimated their long-term importance. I changed my mind in December 2024."
社区会议本周五举行,重点介绍 vLLM 集成
Keras 生态系统最新进展,重点是新版 vLLM 集成,通过 Google Meet 举行,任何人可加入。
Gemini 3.6 Flash ARC-AGI-2 得分 60.4%
Google Gemini 3.6 Flash 在 ARC-AGI-2 验证集上达到 60.4%,每任务成本 $0.61。
LeCun 加入 AI 投资机构 224 Ventures
Yann LeCun 作为合伙人加入 224 Ventures,专注于 AI 初创公司投资。
Google 本可通过开源统领 AI
"Feels like Google could have been the dominating force in AI by open-sourcing Gemini, Veo, and Nano."
Grok 进入 Blender 3D 软件
Elon Musk 发布"Grok in Blender",暗示 AI 模型与 3D 创作工作流的整合。
MiniMax-H3-Turbo-Lora 演示上线 HuggingFace
Turbo LoRA 演示空间已可在 HF 上体验。
Grok Imagine 视频生成持续迭代
xAI 发布 Grok Imagine 新演示,持续扩展 AI 视觉与视频生成能力。

