OpenAI's Jalapeño chip delivers more intelligence from every watt
The first test results for OpenAI's in-house inference chip show a major advance: higher throughput and lower latency in one architecture.
Since announcing Jalapeño, its first custom inference chip, OpenAI has been testing the silicon and the system around it. The company says the results mark a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in a single architecture rather than trading one against the other.
NVIDIA's Vera Rubin NVL72 racks are here — first systems land at Microsoft
Vera Rubin NVL72 has entered production, with manufacturing 100% automated and every tray going together in one minute.
The compute tray is engineered for fast compute, assembly and serviceability. NVIDIA congratulated Microsoft on hosting the first operational Vera Rubin NVL72 system, the opening wave of a rack architecture built for the agentic AI factory.
Stability AI raises $76M with the Big Three labels backing it
Stability AI's new round drew support from Universal, Sony and Warner, making it the first AI company backed by all three major labels. Total funding now stands at $232 million, a strong signal for AI in music and media.
Perplexity launches Portable Computer running fully on-device
Portable Computer is a fully local version of Perplexity Computer, running on NVIDIA DGX Spark. The orchestrator LLM, sub-agent LLM and agent harness all execute on local hardware — no cloud dependency.
"Two labs will soon control most of the world's compute."
Clement Delangue, CEO of Hugging Face
Perplexity's local-agent study: a 27B model scores 82.6%
Perplexity published local-first agent research, where an on-device 27B model scores 82.6% on real knowledge work, beating open-source harnesses Pi and Hermes. Its post-trained PPLX 27B reaches 85.4%.
OpenAI adds $100 Business Premium seats for small teams
ChatGPT now offers Business Premium seats at $100 per month with no five-hour limit and higher usage, aimed at small businesses and startups that want capabilities once reserved for big companies.
Apodex 1.1 releases an open-source agent model and framework
Apodex 1.1 open-sources the FrontierAgent framework and weights, supporting a native command-line TUI, ReAct and Agent Team modes, and it runs locally on macOS and Linux.
Qwen3.8-Flash-Next teaser previews the Qwen4 architecture
The Qwen team launched a pre-release page for Qwen3.8-Flash-Next, described as a Qwen4 architecture preview, with release expected on August 26. Thousands are already watching.
Hot Chips 2026: NVIDIA bets on full-stack agent platform
NVIDIA showed Vera CPU, Vera Rubin, Groq 3 LPX and BlueField-4, arguing extreme co-design is needed for agentic workloads.
NVIDIA unveils Jetson Orin Nano 2
Built for entry-level edge AI and robotics, it brings frontier intelligence to small, power-constrained devices.
Vercel launches Run SDK for safe agent evals
Run SDK executes untrusted code in a hardened QuickJS sandbox, exposing only approved host functions and pausing for human approval.
Vercel Connect is now GA, linking 100+ services
Short-lived scoped access tokens give agents secure access to Slack, Linear, GitHub and more, with RBAC and audit trails.
ExtractBench tests 14 document extraction systems
LlamaIndex evaluated frontier systems on schema-guided extraction across 370 enterprise documents, including scans and nested tables.
SemiAnalysis releases AgentX 1.0 coding benchmark
An open-source multi-turn agent benchmark built from roughly $3 million of real trajectories, running on 1,000+ chips.
Higgsfield offers Ox Alpha free for a limited time
Ox Alpha is temporarily free through Higgsfield Supercomputer, which claims capacity for a quadrillion tokens per day.
ChatGPT can now securely log in and act for users
ChatGPT can securely sign into user accounts and perform actions, extending its agent capabilities into account-driven work.
OpenWorker strengthens security workflows
The open-source agent shipped a new version with security-focused features; Andrew Ng notes it is gaining traction in cybersecurity.
Jensen Huang gifts a DGX Station to Perplexity's CEO
After Arav Srinivas demoed Portable Computer on DGX Spark, Jensen Huang gave him a DGX Station, saying it can serve frontier models like GLM 5.3 locally.
Snowflake's coding agents cut inference costs 33–45%
Snowflake AI Research says CoCo and CoWork cut token spending by up to 66% via context compression, lowering end-to-end trial costs by 33–45%.
Perplexity CEO: agent reasoning should shift to local hardware
Under compute and power constraints, Arav Srinivas argues a large share of agentic inference should run on-device. Portable Computer is that vision in practice.
Qwen3.8-27B enters Code Arena's top 10
Ranked ninth overall, it is the only model of its size class in the top 10.
MiniMax-M3 finishes a real agent task for $0.018
On a real-world agent benchmark, it spun up an inbox and sent a business email end-to-end — the cheapest model on the board.
IBM Granite 4.2 arrives on Ollama
The 3B, 8B and 30B open models are free to use, licensed for commercial work, and tuned for enterprise agents.
SGLang supports Granite 4.2 from day one
A 30B dense reasoning model with native chain-of-thought, AIME25 89.17, and three switchable thinking modes.
MiniMax H3 integration index launches
Awesome MiniMax H3 catalogs everything built around H3, from 24GB-VRAM ComfyUI setups to enterprise SGLang and vLLM-Omni deployments.
Vercel open-sources the minimal coding agent fx
A tiny Zig-based coding agent designed for research and embedding, architected from the outset to avoid software bloat.
ByteDance releases TLive-Omni for live commerce
An omni-modal model that processes images, video, audio and text simultaneously for e-commerce live-streaming.
ByteDance launches the Doubao Work workbench
A standalone office-agent client for work scenarios, with new users getting a free month of standard membership.
OpenAI Build Week names 8 winning projects
Meet the builders behind eight winning projects and what they shipped with Codex.
Adobe Firefly launches a free animation generator
Turn ideas and characters into 2D and 3D animation quickly, aimed at professional creative workflows.
Anthropic open-sources protein binder dataset
A newly generated protein binder dataset is now available on Hugging Face for the research community.
v0 integrates Vercel Connect
Apps and agents connect to Slack, GitHub, Notion and Salesforce without API keys, using short-lived auto-refreshing tokens.
vLLM-Omni adds MiniMax H3 video generation
One model, four tricks: text-to-video, first/last-frame video, lip-synced clips, and green-screen relighting.
Coding agents survive only 5.4% of repo-wide tasks
New research finds agents complete only 5.4% of whole-repository evolution tasks, underscoring current limits.
Autoscientist can auto-configure RL workflows
It now extends to configuring reinforcement-learning recipes, showing the entire post-training stage can be self-improving.
Codex's five-hour limit returns — for Plus only
Codex reintroduced a five-hour usage cap that applies only to Plus users; Pro users are unaffected.
Qwen4's new architecture is coming
Beyond Qwen, many model vendors are moving to entirely new architectures, a notable shift in the open-model field.
ChatGPT extension adds more browsers
Codex can now read your open tabs as context.
DeepSeek-V4-Flash quantized downloads near 20k
A GGUF build is close to 20,000 downloads on Hugging Face.
Cursor expands Grok model quota
Included usage rises as demand grows with Grok 4.6.
Vidu opens its full e-commerce ad API
Five production modes with per-second pricing.
Wan 3.0 ecosystem expands fast
Creators ship 30-second takes and Omni-Reference work.
AI Gateway adds webhook callbacks
Lower idle compute after long generations.
AI Gateway offers zero-config managed tools
Add exa_search with no credentials or keys.
Apodex runs analysis on your local data
Reads local files, runs code, and returns verifiable conclusions.
An AI ad startup pivots to "human ads"
Real-user UGC, no AI — a year after its $12M domain bet.
Agents need designs per firm/worker/agent
Eleven configurations may each require a different design.