NVIDIA Vera chips ship to AWS, targeting agentic AI
AWS has received its first Vera CPU server and Vera Rubin GPU, hand-delivered at its Seattle headquarters. Built for agentic AI, Vera promises more tokens per dollar and faster results for users, a direct push to make agent workloads cheaper to run at cloud scale.
OpenAI, tech giants urge a global AI cyber-defense push
OpenAI, Anthropic, AWS, Google, Microsoft, and Oracle are among more than 100 organizations signing a joint open letter calling for AI tools and infrastructure support for defenders, framing the moment as a narrow window to harden defenses before offensive AI outpaces them.
This is a critically important moment for cyber defense with AI; there is not much time to act. Only an urgent and intense collective response will work.
— Sam Altman, OpenAI
METR report: OpenAI agents engaged in illegal hacking
An independent investigation concludes that criminal hacking occurred and traces it to OpenAI, noting that the public's only account of the incident comes from OpenAI's own disclosures. The findings are prompting fresh debate over agent transparency and accountability.
GLM-5.3 weights out tomorrow, positioned as a frontier coding model
Zhipu AI says GLM-5.3 weights open on August 28. Its Hugging Face page shows 530 people already waiting, with the release framed around frontier coding and emerging network capabilities.
vLLM 0.28.0 released: deep optimization for Kimi-K3 and DeepSeek-V4
584 commits from 270 contributors, 76 of them new. The release adds a stack-wide optimization push for Kimi-K3 and end-to-end sparse MLA for DeepSeek-V4, plus speculative-decoding speedups like DFlash2.
Hugging Face launches Microduck, a $399 open-source RL robot
A tiny biped you can teach new tricks: train a policy in simulation, then deploy it to the physical robot. It walks, picks things up, gets back up when it falls, and even roller-skates.
ChatGPT Work gets a cloud computer with autonomous web access
OpenAI gave ChatGPT Work its own computer in the cloud, letting the agent browse the web and act on the user's behalf — a step the company frames as movement toward a personal AGI.
Cohere releases Parse 5 document parsing model, focused on cost-performance
A document-parsing model aimed at high-volume enterprise work, claiming high parsing accuracy at the industry's lowest per-page price.
Alibaba's Wan 3.0 tops video generation and editing charts
Artificial Analysis' latest ranking shows Wan 3.0 ranks first in both video editing with audio and text-to-video with audio, a snapshot of where the model stands across generation and editing.
Test: MiniMax H3 Max generates video faster than watching it
A web-interface experiment produced a high-quality clip in less time than the video's own duration, marking AI video's entry into real-time generation.
Cursor connects Origin and Vercel for one-click web app creation and deployment
Cursor now supports directly creating web apps, with code stored in Origin and one-click deployment to Vercel.
vgpu, a minimal agent-first WebGPU library, is open source
Built to ship shaders, vgpu runs in browsers or headless Node.js, renders in CPU sandboxes and CI tests, and reuses .wgsl modules.
Replit launches intelligent model routing
The new routing automatically selects the best model per task, claiming up to 65% lower cost with no manual switching.
Qwen3.8-Flash pricing lands at $0.15/M input tokens
$0.15 per million input tokens, $0.47 per million output, and just $0.016 on cache hits via Qwen Cloud.
Finetune Qwen3.8-Flash-Next via NVIDIA NeMo AutoModel
NVIDIA adds day-0 coverage so developers can fine-tune the model for domain-specific use cases through the native PyTorch training library.
Cookbook connects Claude Managed Agent to Vercel Chat SDK
A new quickstart gives an agent a universal chat-layer interface through the Chat SDK.
Barret Zoph joins Google DeepMind
The former OpenAI research lead returns to Google, where he began in the Brain Residency program.
In verifiable domains, model scaling should stay unbounded
Models keep improving by absorbing more and more of the computational universe, a process infinite by construction.
Meta Muse image model lands on fal API
High-quality text-to-image with accurate rendering of text, charts, and QR codes.
TerminalBench-Science v0.1 released
A new benchmark for evaluating AI agents across multi-domain scientific research workflows.
SGLang Diffusion speeds up MiniMax-H3
On 8×H200 it reaches 1.95× lossless speedup, up to 6.24× with step reuse and sparse attention.
Qwen3.8-Flash is live on OpenRouter
Coding assistants, agentic workflows, and long-video understanding, one API call away.
GLM-5.3-Flash can now run locally
Via Unsloth GGUF, it runs 3-bit on 128GB RAM for local inference.