GPT-5.6 Sol gains an Ultrafast mode
OpenAI is previewing an Ultrafast mode that runs GPT-5.6 Sol at up to 14x the speed. It launches first in the OpenAI API to a select group of customers, with access expanding to more businesses as capacity grows. For latency-sensitive agent loops, the leap is the kind that changes what a loop can afford to do.
The frontier 10% run plugins twice as often
OpenAI data shows the top 10% of enterprises use plugins twice as often and skills six times as often as typical firms. The message is blunt: frontier firms are not ahead by accident.
One command to rule them all — tokens, model choice, lower costs, observability, and ZDR. This will become the default way of using coding AI at scale.
Guillermo Rauch, on Vercel's AI Gateway
ChatGPT starts remembering your computer activity
ChatGPT can now remember your activity across the apps and websites on your computer. With Computer History in the desktop app, future interactions feel more personalized and require less explanation.
Claude Code adds an auto-continue checkbox
Hit your usage limit in Claude Code desktop? There is now an auto-continue checkbox that picks up exactly where you left off once your limit resets. Small feature, big difference for long-running agent sessions that get cut off mid-task.
Music3 goes open-weight
MiniMax released Music3, a next-generation, open-weight, production-ready music model. It is already live on ComfyUI, with a cloud release on the way, framed around democratizing AI through open source and open science.
H3 tops the Video Edit Arena
MiniMax claims H3 took the crown on Video Edit Arena — number one not just among open models, but state of the art, full stop. The pitch: drop in your footage, prompt the impossible, and let H3 reimagine every frame.
H3 and Magnific open their showcase
MiniMax's open-source H3 video model, which handles text, image, video, and audio input, is teaming with Magnific for a showcase. The waitlist is open, registered users can watch directly, and works built on H3 and Magnific are being solicited.
Day-zero support for Music3
SGLang ships day-zero support for MiniMax Music3, a music generation model built on a Qwen3 plus RVQ autoregressive backbone with a flow-matching DiT decoder.
Connect coding agents to AI Gateway
A single command connects coding agents to Vercel's AI Gateway, auto-configuring eight popular harnesses with 300+ models from 30+ providers at no markup, plus open-weight models with ZDR and US inference.
Cloud agents start 3x faster
Cursor cloud agents now start three times faster thanks to continuously prepared development environments, letting users hand off ambitious, long-running tasks that execute from start to finish.
Firetiger joins Cursor
The Firetiger team, which builds agents for production software that keep handling issues after deployment, is joining Cursor to close the loop from writing code to operating it.
Grok 4.6 hits the Pareto frontier
Now available in Perplexity and Perplexity Computer; on WANDR it matches Fable 5 results at over 60% lower cost.
Benchmarked as an orchestrator
On Wide-and-Deep-Research, Grok 4.6 sits neatly on the performance-vs-cost frontier for Pro and Max users.
Intelligence per dollar
grok 4.6 impresses on intelligence per dollar; Fable 5 still leads outright, but SpaceX and Cursor are now firmly in the frontier game.
Reinventing the design process for AI
Replit Design guides users with ambient intelligence and one-click variant suggestions, and can create a design system once so everything after automatically follows brand specs and is ready to publish.
API hackathon winners, built in a weekend
Runway named the winners of its API Hackathon, built on Runway Dev — from Quigo, turning passive video into interactive stories, to ClinicalSim, rehearsing hard conversations.
Tokens are the new commodity
NVIDIA argues AI factories are the industrial infrastructure of the AI era and tokens are the new commodity — then asks how to optimize AI token economics.
Gen-4.5 meets Vera Rubin on stage
At the Runway AI Summit, NVIDIA's Richard Kerris showed how real-time generative AI, powered by Vera Rubin, makes creative production a live, artist-controlled process.
SL2T: sign directly to your phone
A sign-language-to-text model built in close collaboration with the Deaf community.
Qwen3.8-27B on the way
An update focused on high "intelligence density," expected on August 14.
Sakana Chat refresh
Powered by Fugu and Namazu, with full code execution and no login required.
eve for persistent agents
A Next.js-style framework for durable AI agents, self-hosted on an open SDK.
SCoPE for video diffusion
A sightline-coordinate positional encoding for video diffusion transformers.
HarnessAgent swaps any brain
Run, swap, or make brains compete — one line of code to switch to Grok.
Could ASI escape watermarking?
Ethan Mollick's verdict on an undetectable AI watermark: the answer is no.
Test-time training, revisited
Francois Chollet on where TTT strongly outperforms: ARC 1-2, for now.
Tool use becomes the bottleneck
Agentic work may soon slow on tool use, not LLM inference speed.
Flash, the workhorse tier
Pichai on shipping Flash updates fast to reach developers' hands.