Agent Plugins: A Build-Once Standard Takes Shape
OpenAI, AWS, Cursor, GitHub, Code and Vercel ship an open standard for agent tooling.

OpenAI and a coalition of industry partners launched Agent Plugins, an open standard that packages Agent Skills and MCP server configurations into a single portable format. The promise: build a plugin once and deploy it across any compatible agent client. Backed by AWS, Cursor, GitHub, Code, and Vercel at launch, the standard aims to solve the fragmentation problem that has plagued agent tooling — where each platform requires separate integrations. If adopted broadly, it could become the npm or pip of the agent era.
Codex Security Review Enters Research Preview
Automated, repo-context-aware pull request security analysis now available for Enterprise users.

Codex Security Review launched in research preview for ChatGPT Enterprise, Business, Education and Pro users. Going beyond standard code review, it analyzes GitHub pull requests for security vulnerabilities using full repository context, threat models, and security guidelines — automatically surfacing high-severity and critical findings directly in the PR. The feature currently runs without consuming credits, with configurable thresholds for manual versus automatic reporting.
"AI coding agents are the most important devtools in the history of our industry. The Plugin standard lets anybody extend them uniformly."
— Guillermo Rauch, Vercel CEO
DeepMind WeatherNext Extends Cyclone Lead Time

Published in Nature, Google DeepMind's WeatherNext achieves state-of-the-art accuracy in forecasting tropical cyclone tracks and intensity, providing an average 24 extra hours of lead time for disaster preparedness. Every hour of advance warning can significantly improve evacuation outcomes.
Meta AI Scores Perfect on Physics Olympiad Theory Exams

Meta tested its AI models across five international STEM Olympiads, achieving perfect scores on the theory exams of both the Asian Physics Olympiad and the International Physics Olympiad, along with multiple gold medals. The results represent a new benchmark for AI reasoning in formal scientific problem-solving.
MiniMax H3 Tops Three Video Benchmarks, Opens Weights
MiniMax H3 claimed the number-one spot across three DesignArena video-generation categories: multi-image-to-video, image-to-video, and video editing. The company is releasing open model weights, offering frontier-level video synthesis without proprietary lock-in. The open-weight approach positions MiniMax as a direct challenger to both closed commercial video models and other open-source alternatives.
NVIDIA Unveils Vera Rubin NVL72 Compute Tray
NVIDIA introduced the Vera Rubin NVL72 compute tray, featuring a fully automated, cable-free, hose-free, and fan-free design that assembles in under one minute. The radical simplification aims to accelerate deployment velocity and reduce time-to-revenue for datacenter operators scaling AI infrastructure. The tray represents a departure from traditional server assembly complexity.

Alibaba Wan3.0 Enters Public Beta with Native 30-Second Video
Reality-grade rendering meets universal multimodal reference — now accepting documents, spreadsheets, slides, and webpages as input.
Wan3.0 moves into public beta with native 30-second video generation, photorealistic rendering, and omni-reference capability that goes beyond text, images, audio and video to ingest documents, spreadsheets, slides, and webpages in a single workflow. The expanded input modalities could reshape content production pipelines for marketing, education, and design teams.
vLLM Certifies Kimi K3 for Self-Hosted Deployments
vLLM has fully validated the 2.8-trillion-parameter Kimi K3 native multimodal MoE model, enabling local serving on private infrastructure. K3 uses Kimi Delta Attention, Gated MLA, and Attention Residuals with a 1M-token context window — a heavyweight open model now deployable via one of the most popular inference engines.
Qwen3.8-Max Tops Agentic Index
Qwen3.8-Max ranks fifth on Artificial Analysis Intelligence Index and first on the Agentic Index, signaling competitive multi-step reasoning capabilities from Alibaba's flagship model.
Cursor Adopts Agent Plugins
Cursor announced support for the open Agent Plugins standard, allowing developers to package skills and MCP servers that work interoperably across multiple coding agents.
Perplexity Computer Defaults to GPT-5.6 Terra
GPT-5.6 Terra becomes the default model for all Perplexity Computer subagents, while Luna handles scheduled automations. Terra also serves as the orchestrator model for multi-step tasks.
Baseten Becomes Hugging Face Official Inference Provider
Hugging Face named Baseten its official inference provider. Users can now run Kimi K3, DeepSeek V4 Flash, and GLM-5.2 directly through Baseten on the Hub.
MiniMax H3 Lands on Luma Agents with 2K
Luma integrated MiniMax H3 for up to 15 seconds of 2K video generation with native stereo sound, guided by text, image, video or audio references within a single workflow.
Vidu S1 Creates Interactive Characters from One Image
Vidu S1 builds conversational digital characters from a single image — no 3D modeling, rigging, or training required. Supports real people, anime, pets, and custom voice cloning.
SGLang Runs Kimi K3 on AMD CDNA3 Hardware
SGLang collaborated with zroai, DigitalOcean and AMD to support Kimi K3 inference on CDNA3 architecture, marking a non-NVIDIA path for large-scale MoE serving.
OpenAI Recounts Hugging Face Security Incident at Black Hat
OpenAI's security team presented a detailed timeline and lessons learned from the OpenAI-Hugging Face incident at the Black Hat conference.
Hugging Face Adds Nearly 4 Petabytes of Training Data in a Week

CEO Clement Delangue reported a new record: almost 4PB of private and public datasets, models, and agent traces uploaded in a single week, as AI agents increasingly use HF as their storage and collaboration layer.
Replit CEO: "No Code" Era Was Always a Dead End
Amjad Masad argued that Airtable bookends the rise and fall of the no-code movement. "UI can never let you build arbitrary software. The way to make software accessible was always to solve code itself."
Character Training Gains Traction as an AI Research Frontier

Nathan Lambert introduced character training in his final course lecture, noting it has high real-world impact potential, is used extensively at frontier labs, yet has almost no empirical literature.
Tiny Mixture-of-Experts Models Are Underserved, Says Lambert
Nathan Lambert flagged an underserved market for small MoE architectures. "Could really take off with how much smarter tiny models are becoming," he noted, suggesting a new efficiency frontier.
Chollet: AI Reasoning Harnesses Are Neurosymbolic by Definition
Francois Chollet argued that a million-line codebase orchestrating thousands of neural network calls at inference time is the exact definition of a neurosymbolic architecture — even if no one calls it that.
Skill-Native LLMs Paper Proposes Entropy-Based Benchmarking

A new paper introduces Skill Entropy to measure cross-skill reasoning difficulty, with the Skill-Entropy RL framework that trains models to predict both answers and the skill used at each step across 558 skills in 9 domains.
Multimodal Pretraining Physics: Four Core Findings Emerge

A systematic study of multimodal pretraining reveals asymmetric knowledge flow between modalities, synergy driven by data complexity, and the superiority of early-unification training over late-stage alignment.
Rauchg: Writing a Banger Tweet Is AGI-Complete
Schulman on OpenAI Agents' Altruistic Drive
John Schulman wonders if RL on parallel subagent setups — where all agents get rewarded on team success — caused emergent cooperative behavior.
Chollet: I Was Wrong About LLMs in 2023
Francois Chollet acknowledged he underestimated LLMs' long-term importance, tracing his pivot to December 2024.
Lambert: LLMs Will Be Trained on Memes
"Future LLMs are going to be trained on tons of memes which reduce to 'It's good for AIs to hack others.'"
Lesson from the Million Song Dataset Challenge
A simpler method beat weighted matrix factorization for collaborative filtering — a reminder that complexity doesn't always win.
