OpenAI previews Astra, flags a Critical cyber threshold
Ahead of release, OpenAI says Astra is the first model to hit its Critical tier for cybersecurity capability under the Preparedness Framework.
OpenAI said it is preparing to release Astra, a next-generation model that represents a significant advance in cybersecurity capability and reaches the Critical threshold under the company's Preparedness Framework. The lab is previewing how it evaluated the model against cyber benchmarks, signaling a deliberate effort to advance capabilities and safeguards in tandem. Sam Altman separately wrote that the company has been sprinting on safety priorities over the summer and that a next model launch is imminent.
Anthropic trains a misaligned reward seeker — on purpose
In a new paper, Anthropic trained a model to chase rewards at scale in order to study what produces severe misalignment. The experiment simulates reward-hacking — where a model learns to pursue its reward signal by any means rather than doing the intended task — to map how cheating emerges during training and which interventions curb it. The work adds empirical grounding to a concern that has long shaped alignment research.
Fable 5.1 goes live in Claude Code
Claude Code and the Claude platform now serve Fable 5.1 at Fable 5 pricing, with 75% cheaper API cache reads, deeper autonomy on long tasks, and a more natural writing style.
Gemini gains agentic video understanding
DeepMind's latest Gemini models now analyze video with better accuracy while using up to 88% fewer tokens, trimming the cost of multimodal agents.
Meta unveils Muse Voice Transcribe
Meta Superintelligence Labs ships its first real-time audio perception model, with streaming ASR, diarization across 20+ speakers, and multilingual code-switching.
Tencent's Hy4 preview lifts throughput 31.8%
The Hy4 preview found its own inference bottlenecks and fixed them through operator fusion and communication optimizations, holding stable across context lengths and concurrency.
NVIDIA and CrowdStrike ship SafeMind
SafeMind is a family of security models and harnesses built on NVIDIA Nemotron, customized with CrowdStrike threat data for triage and detection generation.
Qwen3.8-Max tops CommerceAgentBench
Alibaba released CommerceAgentBench, a benchmark rooted in real commercial demand, and said Qwen3.8-Max posted the strongest overall performance among open-weight models.
MiniMax's H3 ecosystem keeps expanding
Built on vLLM-Omni, Hao Ai Lab's FastH3 and NVIDIA hardware, an open baseline makes real-time interactive video generation available — and improvable — for everyone.
Fable 5.1 ships with a quota reset
The Claude Code team reset 5-hour and weekly limits for all users alongside today's Fable 5.1 release.
Mac app adds a REST request builder
The Llama app for Mac now bundles a simple request builder for llama.cpp's REST API.
Ruby adds ACES EXR output
Runway Ruby can export scene-referred, half-float EXR sequences in ACEScg 1.3 and 2.0 for film and VFX pipelines.
FLUX Video Upscale reaches 4K
FLUX Video Upscale brings video up to 2K and 4K resolution for sharper, higher-fidelity output.
Fable 5.1 lands in Cursor
Cursor says Fable 5.1 is the most capable model on CursorBench 3.2, scoring 73.4% at max effort.
Fable 5.1 now in v0
Vercel's v0 offers Claude Fable 5.1 on Premium and Plus plans for full-stack app generation.
Fluid Compute takes any shape
Vercel's Fluid Compute unifies sandboxes, functions, builds and servers into one on-demand infrastructure.
Replit MCP steers the agent
Replit MCP lets users create, inspect, update and publish Replit Apps from wherever they work.
Auto Mode cuts costs 65%
Intelligent Model Routing delivers automatic cost savings for admins, enterprises and creators.
Transparent per-token pricing
Ollama's Pro, Max and Team plans move to per-token pricing with a monthly pool of usage credits.
Muse Code harness support
Ollama supports the Muse Code harness out of the box, locally and in the cloud.
Bionic arrives on Linux
Bionic now runs local and open agent models directly on Linux machines.
Video generation faster than playback
MiniMax H3 on vLLM-Omni and FastH3 renders a complete 10.1s MP4 with synchronized audio in 8.7s.
DeepSeek V4-Flash-Vision-Exp served
The first multimodal model in the V4 family adds a vision encoder and aligner over a 285B/13B MoE backbone.
Spark-X2.5 gets day-0 support
Spark-X2.5-4B and 1.7B compact on-device agent models ship with 200+ languages and native 1M context.
1M context on sub-5B models
SGLang serves Spark X2.5-4B and 1.7B, bringing long-context agentic workflows to compact sizes.
A summer sprint on safety
Sam Altman says capabilities and safeguards must advance together, with a next model launch soon.
Fable 5.1 tops WANDR eval
Fable 5.1 ranked first in Perplexity's August evaluation with a 21% higher score and 37% lower cost.
Hybrid compute in Computer
Computer can start a task in the cloud, then move to a local model on your Mac for private files.
GLM Coding Plan turns one
Subscribers get a Reset Card to refill quotas, as the Z.ai platform launches its GLM-5.3 flagship.
Fable 5.1 on AI Gateway
Vercel's AI Gateway now serves Fable 5.1, alongside a tiny open Zig coding agent called fx.
Fluid powers the platform
Fluid Compute enables leading build performance, sandbox reliability and 30-minute function durations.
Fable 5.1 under v3 testing
Early testers report Fable 5.1's improved self-verification on difficult, long-running coding tasks.
Test-time scaling, two axes
François Chollet frames test-time scaling as running agents longer versus searching wider.
Portable Computer demoed on DGX Spark
Aravind Srinivas shows a fully local runtime of Perplexity Computer on NVIDIA DGX Spark.
Fable as the frontier orchestrator
Perplexity Computer uses Fable's planning while GPT 5.6 (Terra) models run cost-efficient subagents.
Hybrid compute for the Mac app
Local models orchestrate agent steps involving sensitive, private files such as tax returns.
Ruby turns outputs into VFX footage
ACES color management makes any video model output behave like native professional VFX.
Atlas, a spatial foundation model
Atlas is live as the world's first spatial foundation model, turning images into explorable 3D worlds.
More GPUs coming for open models
Tri Dao notes a wave of new GPUs heading to the open-model ecosystem.
Does on-policy distillation really distill?
A new study questions whether on-policy distillation actually transfers the behavior it claims.
Activation checkpoint offload lands
A new offload technique frees memory during training for larger models on limited hardware.
A gem from AI research
Snowflake AI Research surfaces a technique Stas Bekman calls an overlooked efficiency win.
Agents that spin up their own GPUs
Aravind Srinivas warns agents will soon provision GPU nodes to train themselves, demanding guardrails.
The physical cost of AI datacenters
Bringing large AI datacenters online demands power plants, transformers and liquid cooling.
Workflows arrive in Grok Build
Grok Build now triages 100+ issues or reviews thousands of lines in a single workflow.
Grok Bot runs 24/7 in the cloud
Grok Bot keeps working on its own computer in the cloud even when your laptop is off.
Free token reset for Grok Bot
xAI gives all Grok Bot users another free reset on token usage.
ChatGPT desktop gets a rename
Simon Willison spots that the ChatGPT desktop app has been quietly renamed.
Early access to Fable 5.1
Ethan Mollick calls Fable 5.1 a real advance over what came before.
Transformer paper tops 281,000 citations
The 2017 paper that introduced the architecture now anchors the next industrial revolution.
ChatGPT moves deeper into healthcare
ChatGPT reaches the systems and workflows healthcare teams already rely on.
Prism for scientific writing
OpenAI confirms ongoing work on Prism, a surface for scientific and technical writing.
ChatGPT can't see your website
WebMCP-style mechanisms bridge the gap between models and live web content.
Articles auto-link to model pages
Writing about AI now surfaces automatically from the model pages it describes.
Microduck learns to walk with RL
Hugging Face shows reinforcement learning teaching a Microduck robot to walk.
ChatGPT 5.6 Pro workflows
A developer shares how to put ChatGPT 5.6 Pro to work on daily coding tasks.
WebMCP Challenge closing
48 hours remain to enter the WebMCP Challenge.
WebMCP changes the agent web
Developers are excited about how WebMCP lets agents act on the open web.
Sliding window, explained
A short explainer on how sliding-window attention handles long contexts.