DeepSeek V4-Flash Official API Launches, Agent Capabilities Surge
The official API enters public beta with Agent performance far surpassing V4-Pro-Preview and native MCP support built in.

DeepSeek has officially launched the V4-Flash API in public beta, marking a major milestone for the company's lightweight model strategy. The new release delivers a substantial leap in Agent capabilities, with benchmark scores now far exceeding the V4-Pro-Preview across multiple evaluation dimensions. The official V4-Flash API natively supports the Model Context Protocol, enabling seamless integration with developer toolchains and autonomous agent workflows. The announcement drew over five million views and twenty-two thousand likes in the first twenty-four hours, reflecting intense developer demand for cost-efficient, high-performance agent infrastructure. Engineers can now access the endpoint directly, with the team promising further capability upgrades in subsequent releases.
Gemini Robotics 2: 20 Minutes of Uninterrupted Autonomy
DeepMind demonstrates Gemini Robotics 2 performing continuous real-time tool manipulation on the FR3 Duo platform.
Google DeepMind has unveiled Gemini Robotics 2, a next-generation robotics foundation model demonstrated on the FR3 Duo robot. The system performed twenty minutes of uninterrupted, real-time tool manipulation without any human intervention. This represents a significant advance beyond short demonstration clips, showcasing sustained dexterity, environmental adaptation, and integrated vision-language-action reasoning within a single unified architecture. Industry observers view the twenty-minute milestone as a credible step toward practical, long-horizon deployment in manufacturing, logistics, and laboratory automation. DeepMind has not yet announced a public release timeline for the model.
OpenAI to Retire GPT-5.4 Series from ChatGPT Login
OpenAI announced that GPT-5.4 and GPT-5.4 mini will no longer be available to users signed in with ChatGPT starting August 31. The models will remain accessible through the OpenAI API and Codex sessions authenticated with an API key. Users are advised to update workspace defaults, saved model configurations, custom agents, and scheduled tasks before the cutoff date. The company also highlighted desktop application updates including multi-folder local project support, code browsing and review tools, and enhanced image editing capabilities.

Qwen-Audio-3.0-ASR-Flash Released
Alibaba's Qwen team has shipped Qwen-Audio-3.0-ASR-Flash, an upgraded automatic speech recognition model supporting context consistency, domain terminology recognition, custom hotwords, and speech-to-structured-transcript polishing. The Flash variant emphasizes low-latency inference for real-time transcription, with internal tests showing strong performance on specialized medical and legal vocabulary.
MiniMax H3 Launch: Omni-Reference, Open Weights
MiniMax has officially released the H3 model with Omni-Reference capability for commercial-grade generation, emphasizing cost efficiency and open-weight availability. H3 is now available on the MiniMax API, while the Hailuo AI product supports video and image generation from text or photos. Multiple platform partners including Runway, Vercel AI Gateway, Pika, Krea, OpenRouter, fal, and Leonardo AI integrated H3 on day zero, signaling broad industry adoption of the model as a creative infrastructure layer.
Omni-Reference capability means the model takes your entire creative vision, not just a single text prompt.
@MiniMax_AI on H3 launch

Hy-MT2 Hits 700K Downloads, Tops Hugging Face
Tencent's Hy-MT2 open-source translation model has reached over 700,000 downloads since its May release, with the 1.8B variant hitting number one on Hugging Face trending and the 30B-A3B reaching number four. More than 70 verified product and project integrations are now live, with growing ecosystem support across Apple platforms.

NVIDIA Riva Cuts StudyFetch Inference Costs by Nearly 10x
StudyFetch reduced its largest AI inference workload cost by nearly ten times using NVIDIA Riva, Parakeet ASR, and NVIDIA NIM microservices. The efficiency gains support voice tutoring, real-time personalization, and a new agentic learning platform.

Cohere Signs EU Code of Practice on AI Transparency
Cohere has become one of the first companies worldwide to sign the EU Code of Practice on Transparency of AI-Generated Content. The company argues that AI is useless if users cannot trust it, and that building differently requires prioritizing verifiability and clear labeling of AI-generated outputs.
Seedance 2.5 Coming to Luma
Luma Labs teased Seedance 2.5 on its platform, inviting creators to be there when the latest video generation model lands.
Replit Design Goes Live, Model Selector Ships
Replit launched its biggest design update alongside a Model Selector for choosing optimal intelligence per task, Follow-up Tasks for autonomous planning, and a rebuilt usage page.
Replit Design Full Walkthrough Released
A new video covers templates for fast starts, ambient intelligence so builders are never stuck, and design systems ensuring visual coherence across projects.