DeepSeek Puts Vision Into Its Flash Line
A new experimental multimodal model matches V4-Flash on text and makes a major leap on multimodal agent benchmarks — at Flash prices.
After weeks of anticipation, DeepSeek has quietly opened its first multimodal model to the public API. DeepSeek-V4-Flash-Vision-Exp keeps the text capabilities of the V4-Flash line — agents, reasoning and world knowledge — while adding vision tuned for multimodal agent workloads. On multimodal agent benchmarks, the team reports a major jump over the text-only baseline, and it does so at the same price point as V4-Flash.
The model is explicitly experimental, and DeepSeek has not yet committed to open weights. But the message is clear: the multimodal frontier is now being sold at Flash economics, and the lab that made cheap reasoning famous is extending the same playbook to seeing.
OpenAI Cuts GPT-5.6 Sol Pricing by More Than 20%
API and credit prices drop for the next three months as efficiency improves.
OpenAI is making its frontier cheaper again. API and credit pricing for GPT-5.6 Sol falls by over 20% for the next three months as the company pushes the frontier of efficiency. For builders, token-based Codex plans stretch further while the usage included in a subscription stays unchanged — a nudge to ship on the new model.
Grok 4.6 Lands on Google Vertex AI
xAI's flagship spreads beyond the X ecosystem with a 500k context window.
xAI's flagship is spreading beyond the X ecosystem. Grok 4.6 is now served through Google Cloud Vertex AI, with a 500k context window and adjustable reasoning effort from low to ultra-high. Pricing runs $2 per million input tokens, $0.5 for cached input and $6 for output — a serious bid for long-running agents and vision workloads.
NVIDIA's Coding Agent Scores 100% on ARC-AGI-3
NVIDIA's general-purpose coding agent, AVO, posted a perfect score on the ARC-AGI-3 interactive reasoning benchmark — the first flawless result to make the rounds in AI circles this month. The run closes several long-standing gaps and puts a fresh data point on how far agentic coding has come.
"This is very nice work from NVIDIA. Like all high-performing approaches on ARC-AGI-3, it uses deep-learning-guided on-the-fly synthesis of symbolic world models."
— François Chollet, creator of ARC-AGI
Remote Control Gets a Reliability Overhaul
ClaudeDevs spent the month on the feature users asked to fix most, shipping a tighter, more dependable Remote Control.
Track Spend by API Key, Set Hard Limits
Usage and Spend dashboards now break down cost per key, with monthly org and project limits that stop traffic when reached.
Kimi K3 Arrives on Pro and Max Plans
Kimi K3 is now available through included usage on all Ollama Pro and Max subscriptions, with more models promised.
Tenet: A Legal Model Post-Trained on Kimi K3
Harvey introduces Tenet, its first model post-trained for legal work on a Kimi K3 base, built with Fireworks AI.
100M Free GLM-5.3 Tokens for New Users
Build Week becomes an ongoing series, with 50,000 new ZCode users each getting 100M free GLM-5.3 tokens through Aug 23.
Codex Hits 20 Million Weekly Users
OpenAI's coding agent passes a 20M weekly active milestone and hands every user a custom reset allowance.
Runway Ruby Converts SDR Video to 16-bit HDR
Runway's new model Ruby converts SDR footage up to 16-bit HDR, outputting ProRes and EXR sequences compatible with any existing upload or generation up to 30 seconds. For Max and Enterprise plans, it means cinema-grade color and dynamic range without re-shooting.
NVIDIA Vera Rubin Ramps Into Full Production
NVIDIA's next-generation Vera Rubin platform is ramping into full production, a milestone the company credited to its Microsoft partnership.
100,000 Hours of Human Hands, Open and Annotated
Hugging Face and Lightwheel open-source EgoSuite-Open100K — 100,000 hours of egocentric human work in LeRobot format, the largest fully annotated dataset of its kind.
SkyRL's IsoExec Aligns vLLM and Megatron Logprobs
Floating-point non-associativity can make a rollout engine and a trainer disagree on a token's logprob. IsoExec's execution contract unifies them for cleaner RL training.
Ling-3.0-flash-dspark Chases the Batch-1 Floor
AntLingAGI's new model, trained and served with SGLang, targets the low-latency Batch-1 speculative decode on Blackwell.
Pika Speech: A 3B TTS Model at RTF 0.02
Pika's 3B text-to-speech model generates a minute of 48 kHz studio-quality speech in about 1.2 seconds and supports requests up to five minutes.
NVFP4 + DFlash2 Recipes Hit the SGLang Cookbook
Alibaba's Qwen team drops fresh SGLang recipes for Qwen3.8-27B, pairing NVFP4 and DFlash2 for faster serving.
Tencent's Hy-MT2 Models Land on OpenRouter
Hy-MT2-1.8B and Hy-MT2-30B-A3B bring 33-language translation with strong instruction following to the OpenRouter catalog.
v0 Apps Connect to 100+ Services Securely
Apps built in v0 can now reach Slack, GitHub and Salesforce through Vercel Connect, using team connectors and short-lived tokens for auth.
"Is Agentic" Scores a Site's Readiness for AI Agents
Vercel's tool grades how discoverable and usable a URL is for agents, checking rendering, HTTP behavior, error recovery and controls.
fx: A Zig Coding Agent That Fits on Two Floppy Disks
The tiny, open, native agent shrinks again in 0.0.5, xz-compressed to floppy size, and ships tomorrow with its most-requested feature.
Touch Is Robotics' Most Under-Explored Sense
Jim Fan argues touch is criminally neglected: "Imagine sleight of hand wearing thick oven mitts. That's how a robot feels today."
Which Video Model Keeps a Voice Consistent?
Glif tested 20 video models on the same character across new scenes to find which one holds the voice steady.
Anthropic's Claude Academy Is Now Live
The free, no-sign-in course library for Claude is open to everyone, extending Anthropic's run of hands-on AI education.
"It's going to be an era of contradictions: everyone will say they hate AI, and everyone will secretly use it all the time."
— Ethan Mollick
vLLM Conference Kicks Off Next Week
Ray, Google Cloud and AMD happy hours anchor a packed schedule for the first vLLM Conference.
One-Shot Learning Comes to Robotics
Show the robot once and it learns — the pattern that transformed LLMs is now arriving for embodied systems.
Anthropic's ELI5 Skill Is One Prompt
The explain-it-like-I'm-five skill renders plain-language HTML and turns out to be a single well-written prompt.
Follow the Gradient of Surprise
François Chollet's short recipe for research: chase whatever genuinely surprises you.
Hard Limits Stop Traffic When Reached
Organization and project spend caps now include hard limits that cut off usage the moment they're hit.
Sol Credits Go Further in Codex
Token-based Codex plans benefit most from the 20% price cut, while subscription usage stays the same.
Build Week Becomes a Series
Community projects with ZCode and GLM-5.3 earn the program an open-ended extension.
Tenet Targets Legal Work
The first legal post-trained model built on a Kimi K3 base enters the arena.