Infrastructure · OpenAI
Habitat, the storage behind ChatGPT, grows tenfold in a year
Habitat is OpenAI's online storage platform, and the company says it has grown more than 10x year over year. Before its Rust rewrite, the Python service behind it handled more than 20 million requests per second at peak — a reminder that the most visible AI products sit on top of unglamorous, hard-won infrastructure.
Model Release · OpenAI
GPT-Live-1 launches; Yelp answers reservations
GPT-Live-1 listens while it speaks, so callers can interrupt, add details, or change direction mid-conversation. Yelp is using it to make restaurant reservation calls feel more natural.
Engineering · OpenAI
For GPT-6 Astra, less scaffolding is more
OpenAI's guidance for GPT-6 Astra is a cleanup job: revisit skills, AGENTS.md, and task prompts. Make skill triggers specific, load guidance only when relevant, and define what "done" looks like. As models grow stronger, verbose old scaffolding can conflict, so OpenAI argues for shorter skill descriptions, progressive disclosure, and fewer mandatory rules.
Tooling · Anthropic
Claude Code adds plugin eval
A new command, claude plugin eval, lets teams quantify what a plugin actually adds. You create test cases, run a plugin or skill against them, score the runs, then run each case again without the plugin to compare the difference — a quick way to tell whether an extension is doing real work or just adding noise.
Commentary · Mathematics
Fields Medalists warn AI benchmarks are misaligned with math
Twenty-five Fields Medalists, including Terence Tao, issued a statement warning that AI companies advancing by using solved math problems as a benchmark are badly misaligned with the goals of the mathematical community. Mathematics is about conceptual understanding and insight, they argue, not just answers. Mass-produced true-or-false conclusions, rushed publication, and weak citation of prior work threaten the field — and the same tension is spreading to other sciences and creative fields.
Commentary · Hugging Face
On AI extinction risk, listen to the right experts
Hugging Face CEO Clément Delangue pushed back on the way AI extinction risk is debated: asking someone outside the field is like asking your AC technician about climate change. The point is not that the concern is wrong, but that it deserves the full range of expertise from inside the ecosystem.
"I don't want to become a mathematician anymore."
A sentiment François Chollet hears often among math students — alongside his prediction that the digital artists who came of age around 2022 may be the last generation.
Model Release · Sakana
Sakana ships Fugu Ultra v2 and Fugu Max
Sakana AI released Fugu Ultra v2 and Fugu Max. Ultra v2 claims to surpass GPT-6 Astra and Claude Fable 5.1 on some benchmarks such as DeepSWE, while Max targets price-performance — adding models like Nemotron and pricing output 40–60% below GPT-5.6 Terra, Claude Sonnet 5, and Kimi K3. Both are available the same day through an OpenAI-compatible API.
DeepSeek-V4.1-Flash lands on Ollama Cloud
Hosted in the US and Europe with zero data retention; prompts and responses are never logged or trained on. Per-token pricing matches the DeepSeek API, including off-peak rates.
SGLang hits 873 tok/s on DeepSeek V4.1 Flash
Within 24 hours of release, the team pushed throughput to 873 tok/s at batch size 1 on four GB300s using FP8 GEMM fast paths, kernel fusion, DSpark optimization, and MoE TP4.
AMD and vLLM speed up MiniMax M3 by 3–4x
A new post walks through the optimization path for MiniMax M3 on Instinct MI355X — starting from the bottleneck to reach 3–4x serving gains, with tuning advice you can reuse.
Claude Tag triages on-call alerts
When an alert fires in Slack, Claude pulls metrics, diffs deploys, checks flags, finds a likely cause, and proposes a fix teams can approve and merge.
Pose-control models rebuild a memory never filmed
Restored archival photos paired with pose-control models captured the mannerisms of a couple's first meeting for the documentary "Love, Rendered."
Codex desktop can now call Ollama models
After upgrading to Ollama 0.34, ChatGPT Desktop — the Codex app — can be configured to use local Ollama models.
vLLM v0.29.0 makes Model Runner V2 the default
The release spans 594 commits from 277 contributors, including optimizations for Kimi-K3 latency tails and DeepSeek-V4 shared-expert fusion.
MiniMax Design upgrades with Astra and MCP
GPT-6 Astra is live, with new MCP integrations for Blender, Photoshop, and AE, plus collaborative projects and one-prompt 3D generation.
Runway now works inside ChatGPT via Astra
A demo goes from a single product image to a finished animation: Runway builds the style frame, Blender handles the animation, and Seedance 2.5 produces the final render.
Tailscale's model router runs on Vercel AI Gateway
Aperture gives any user in a tailnet instant access to hundreds of models based on identity, with zero data retention, zero markup, free BYOK, and cost-and-usage data on every response.
Replit–Databricks integration reaches GA
Apps built with Replit Agent can deploy straight back to Databricks, inheriting existing security and governance, now with native Lakebase support.
Astra retimes traffic lights with real camera data
Higgsfield used real traffic-camera data to simulate different light timings and rebuilt the streets in Blender to visualize how each change affects flow.
GPT Image 2.5 vs. Nano Banana Pro on 360° consistency
Higgsfield compares the two image models on generating consistent 360° character and room sequences from a single reference image per scene.
Synthesia lets you build custom avatars from a prompt
Realistic presenters, branded mascots, or stylized characters can now be generated from a text prompt or fine-tuned panel controls, replacing stock figures.
Cognition launches SWE-2, with SGLang's SpecForge involved
SGLang congratulated Cognition on the launch, noting SpecForge helped improve speculative-decoding acceptance rates within the training stack.
Solaris: a world model for interfaces
A new demo shows a UI world model that generates interactive visuals in real time and responds to user actions.
Replit buys a company built entirely on Replit
CEO Amjad Masad announced the acquisition and predicted it is only the first case of an AI-native company being bought.
Andrew Ng: AI engineering decides what gets built
Ng argues that AI engineering skills let people influence what to build and drive the build loop, and shares a list of the key skills.
Datasette ships security releases after an AI audit
Two security updates follow an audit run with Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra; public instances using auth plugins should upgrade.
Forecasters peg AI catastrophe odds at 0.47% by 2030
An automated forecasting system from leading experts estimates a 0.47% chance of an AI-generated mass catastrophe by 2030, and a 1.1% chance of catastrophe from any cause.
Importance sampling, the workhorse of RL
A thread explains importance sampling as it appears throughout RL research — the PPO objective, the training-versus-inference mismatch, and more.
MiniMax joins the Nebius AI Builder program
Builders in the program can call MiniMax alongside models and tools from NVIDIA, LangChain, Hugging Face, Cognition, and others, and receive credits and support.
LlamaParse upgrades page-level confidence
The new high-intensity mode provides page-level confidence scores with explanations and lets users refer back to the original document.
Recraft Studio integrates Gemini Omni Flash 1.1
Gemini Omni Flash 1.1 is now available in Recraft Studio, suited for generating new video assets by combining up to five reference images.
GPT-6 Astra Challenge opens submissions
OpenAI opened submissions for the GPT-6 Astra Challenge on Product Hunt, with a deadline of September 18.
The benchmark curve is running out of room
With METR long-horizon tasks effectively saturated — pre-Fable agents did 18 weeks of human work — Mollick asks which graph still shows the exponential climb as Frontier Math and ARC-AGI also top out.
Anthropic report renews data-visibility debate
A post argues the report confirms Anthropic can see user information far beyond chat records, and disclosed data it considered possibly illegal without any law-enforcement request.
OpenAI pauses new Pro signups
OpenAI has paused new Pro subscriptions, leaving existing subscribers unaffected, and gave no timetable beyond "when resources allow."
The Pragmatic Engineer interviews Codex's Tibo
The episode covers Codex's origins, technical decisions, engineering culture, and how software development itself is changing.
IFM_AI's K2 Horizon team heads to CMU
Director Hector Liu will speak at CMU; IFM_AI trained K2 Horizon, currently the strongest fully open language model.