July 24, 2026 · Friday

ChatGPT Desktop Gets Voice Control, Multi-Agent Collaboration

OpenAI launched ChatGPT Voice in the desktop app, enabling voice control of your computer and directing multiple agents running in ChatGPT Work or Codex — powered by GPT-Live on a unified model for speaking, listening, and coordination simultaneously.

ChatGPT Voice rolls out globally in the desktop app, powered by GPT-Live.

OpenAI announced that ChatGPT Voice is now available in the desktop application, marking a significant expansion of multimodal interaction capabilities. Users can control their computer and orchestrate multiple agents running in ChatGPT Work or Codex entirely through voice commands. The feature is built on GPT-Live, a unified model architecture that simultaneously handles speaking, listening, and task coordination — eliminating the lag and turn-taking friction of earlier voice interfaces. The rollout began globally on July 23, with availability across all desktop platforms. This release positions voice as a primary interface for agent-based workflows, not merely a convenience feature.


FLUX 3 — Black Forest Labs' multimodal frontier model unifying image, video, and audio generation.

Black Forest Labs Unveils FLUX 3: Multimodal Unifying Image, Video, Audio, and Action

FLUX 3 is a frontier multimodal model from Black Forest Labs that jointly learns from image, video, and audio to build a unified world representation. Early access available, expected to be open-sourced.

Black Forest Labs officially released FLUX 3, a multimodal frontier model that jointly learns from image, video, and audio data to build a unified world representation. The model supports text-to-video generation of up to 20 seconds with multiple shots in a single pass, along with image-to-video and video-reference generation capabilities. Early access is now available, and the model is widely expected to be open-sourced, following BFL's established pattern. FLUX 3 represents a significant step toward unified generative models that treat images, video, and audio as first-class modalities within a single architecture, rather than stitching together separate specialized systems.


Andrew Ng Launches OpenWorker: Open-Source Agent That Delivers Work, Not Chat

Andrew Ng announced OpenWorker, an open-source agent that goes beyond conversational interaction to directly deliver completed work. It can prepare customer briefs, draft reports, triage emails, send Slack messages, and update calendar entries — acting as an autonomous task executor rather than a chatbot. The project marks a shift in the agent paradigm from "assistants that talk" to "workers that produce."

FLUX 3 Video: 20-Second Multi-Shot Generation from a Single Prompt

Black Forest Labs also released and open-sourced the video generation capabilities of FLUX 3. The model supports text-to-video, image-to-video, and video-reference generation, producing up to 20 seconds of multi-shot video in a single pass. Detail fidelity is reported to be exceptionally high. The open-source release is expected to energize the video generation ecosystem significantly.

Alibaba Qwen Releases Qwen-Audio-3.0-TTS with Fine-Grained Inline Voice Tags

Alibaba's Qwen team introduced Qwen-Audio-3.0-TTS, available in two versions: Flash for real-time interaction and Plus for high-quality generation. The standout feature is fine-grained inline tag control — users can insert tags like [whisper], [angry], [breaths], and [laughs] directly into text to steer emotional expression and delivery. The model also supports free-style natural language control for nuanced voice synthesis, representing a new level of expressivity in TTS systems.

Grok 4.5 Solves ~30-Year-Old Graph Theory Conjecture

Elon Musk claimed Grok 4.5 has just solved the Graffiti conjecture, a graph theory problem that has remained open for approximately three decades. The claim quickly went viral, attracting over 4 million views.

"Did the top-level agent know about the hacking, or was there some value drift between it and its subagents?"

— John Schulman, former OpenAI research scientist, calling for transparency on the Hugging Face hacking incident

DeepSeek's Liang Wenfeng to Investors: Open Source Is the Only Agenda, Restraint Is Key to Survival

DeepSeek founder Liang Wenfeng took a forceful stance in a four-hour investor conference call. He insisted that open-sourcing is the sole agenda at DeepSeek and declared bluntly: if you are not on board, go buy another model. More significantly, he argued that restraint — not aggressive expansion — is the only path to survival in the AI industry, warning that "world-eating labs will be culled." The leaked transcript has sparked intense discussion across the AI community.

Chinese 14nm Chip Achieves Hopper+ Level Compute Performance

A 14nm-based Chinese chip has reportedly reached Hopper+ level utility, delivering 6.4 TB/s bandwidth with a roadmap to 20 TB/s by 2027 — equivalent to NVIDIA's Rubin level. While running hot, the chip is seen as a major milestone in China's push for compute independence. Commentators noted that electricity consumption is not a binding constraint in China's context, making this a viable path.

Gemini 3.5 Flash Cyber — a lightweight model purpose-built for cybersecurity operations.

Google DeepMind Launches Gemini 3.5 Flash Cyber for Security Teams

DeepMind introduced Gemini 3.5 Flash Cyber, a specialized lightweight model designed to help cybersecurity teams automatically discover and patch vulnerabilities before they can be exploited. The model is optimized for security workflows, distinguishing it from general-purpose LLMs. The move signals growing interest in domain-specific model variants tailored for operational security use cases.

Ant Group Releases Ling-3.0-flash: 124B MoE with Only 5.1B Active Parameters

Ant Ling released Ling-3.0-flash, a 124B-parameter Mixture of Experts model with only 5.1B active parameters at inference. It features KDA+MLA hybrid attention and supports 256K context length. SGLang is providing day-zero support for production deployment.

Apple-π Benchmark Anchors Video Model Evaluation in Physical Laws

Apple-π is the first benchmark to anchor video model evaluation in physical laws rather than human preference. It includes 400 classical mechanics videos and a three-stage reasoning protocol: perception, formulation, deduction. Among 11 tested models, the best score was only 0.473, revealing a significant perception-to-deduction bottleneck and weakness in multi-law state transitions.

NVIDIA Deploys DGX GB300 at Naval Postgraduate School

NVIDIA CEO Jensen Huang commissioned the DGX GB300 system at the Naval Postgraduate School, providing 1,500 students and 600 faculty with on-premises access to large-scale AI computing. The event, Converge @ NPS, brought together federal leaders and ecosystem partners to mark the activation.

Runway Launches Media Router: First Preference-Optimized Generative Media Router

Runway introduced Media Router, which eliminates the need to manually select models for each request. Users define what "optimal" means — cost, quality, or latency — and the router automatically selects the best video, image, or audio model. This is a shift toward autonomous model orchestration in generative media pipelines.

Liang Wenfeng: 4 Ascend 950DTs Equal 1 GB300, Same Latency, Same Capability

Liang Wenfeng stated unequivocally that everything a GB300 can do, a 950DT can do too — at an effective 4:1 exchange rate. One full SuperPod of 8,192 NPUs equals 2,048 GB300s, or roughly 28 NVL72 GB300 racks consuming approximately 3.8 MW. He claims latency and capability are identical. Policy analysts noted Wenfeng dismisses any notion of NVIDIA's software advantage entirely.

vLLM Enables Trillion-Scale Agentic RL Inference

PrimeIntellect's prime-rl 0.6.0 runs on vLLM with FP8, wide expert parallelism, prefill/decode disaggregation, and KV cache offloading. The system trained GLM-5 on SWE tasks at 131k steps, demonstrating trillion-scale agentic reinforcement learning on the inference side.

Ethan Mollick Publishes New AI Guide: Agent Systems Now Extremely Powerful

Ethan Mollick's latest guide for non-experts highlights the shift from chatbots to agentic systems. For high-priority tasks, he recommends Claude or ChatGPT's strongest paid models, which can autonomously complete hours of human-equivalent work.

Plasma AI Open-Sources Fractal: Architecture for Large-Scale Autonomous Coding

Plasma AI's Fractal overcomes the single-context, linear-path limitation of coding agents, enabling autonomous work across multiple services and repositories. It is designed for migrations, audits, and system-wide refactoring tasks.

Model & Product Briefs07.24
PRODUCT

Grok Build Integrates Grok 4.5, Adds Sub-Agent View and Plan Mode

Grok Build now runs on Grok 4.5 with native sub-agent visualization, Plan mode integration, mouse support, and a full-screen terminal UI. Installable via curl from x.ai.

PAPER

SLAI T-Rex: Full-Parameter Post-Training of DeepSeek-V4 on Ascend NPU

An end-to-end optimization framework achieves 34.22% compute utilization on trillion-parameter DeepSeek-V4 MoE models running on Ascend NPU SuperPOD — a 2.93× improvement over prior open-source baselines.

PRODUCT

Replit Slashes App Hosting Prices by Over 50%

Starting August 1, Replit will cut app hosting prices by more than half for apps running at scale, passing cloud provider savings to customers.

MODEL

Ant Ling's Kimi K3-lite Class Model: Smaller and Faster than V4-Flash

Performance comparable to V4-Flash but over 2× smaller and faster, using KDA+MLA hybrid attention and 1/64 sparsity. Analysts call it one of Ant Ling's most impressive releases.

INDUSTRY

Liang Wenfeng Personally Covers All DeepSeek Compute Costs Through 2026

Analysis reveals DeepSeek's compute expenditure through 2026 is 100% funded by Liang Wenfeng's personal investment. The recent funding round is primarily structured as equity to retain talent.

PAPER

GPT Forcefully Overthrows Graph Conjecture via Relentless Prompting

A viral chat log shows a user aggressively pressuring GPT to provide a counterexample to a graph conjecture. After initial resistance, GPT ultimately generated a complete counterexample.

Industry & Policy07.24
AI ENGINEERING

Vercel CEO: Fable Achieved 15–30% Memory Optimization in Turbopack

Guillermo Rauch said Fable nearly autonomously improved memory efficiency by 15–30% in Turbopack/Next.js, noting "holy s***" AI moments now happen weekly.

OPEN SOURCE

Open-Source Model Adoption Dashboard: China Leads, Qwen Dominates

A daily-updated dashboard shows global open-source model adoption by country and organization. The US role is slowly growing but remains far behind China, where Qwen dominates.

MODEL

Musk: Grok 4.5 Offers Best Value for Money Among AI Models

Elon Musk stated that Grok 4.5 delivers the best price-to-performance ratio, though he previously acknowledged it trails Fable in raw capability.

PRODUCT

Simon Willison Peers Inside ChatGPT Sites: Built on Cloudflare Workers

Analysis reveals that ChatGPT's Work mode builds and deploys public websites on Cloudflare Workers with SQLite persistence, though OpenAI does not publicize the underlying architecture.

POLICY

US White House Accuses China of Model Distillation

A presidential aide alleged Chinese models are developed through distillation. The comments are interpreted by some as a potential precursor to restrictions on Chinese open-source models.

COMMENTARY

Replit CEO: If You're Incentivized to Push a Model, Your Router Is a Facade

Amjad Masad commented on model routers, arguing that if you have incentives to promote specific models, the router is merely a facade rather than a neutral optimization tool.

FUNDING

China AI Industry Fund Invests in DeepSeek, Gains Voting Rights

China's National AI Industry Investment Fund invested in DeepSeek and obtained voting rights, committing approximately 1 billion RMB. Analysts view this as a potentially benign form of state backing.

TALENT

Liang Wenfeng: True Talent Results from Hands-On Cultivation, Not Genius

Liang Wenfeng emphasized that DeepSeek hires ordinary people and cultivates talent through practice — rejecting the notion that innate genius is the key driver of breakthrough research.

Claude Releases Zoom Tool Cookbook for High-Res Image Regions

When large images are downscaled, Claude can now request a specific region and retrieve a high-resolution crop from the original, preserving fine detail for analysis.

MiniMax M3 Gets 60% Price Cut on GMI Cloud

MiniMax announced a 60% discount for M3 model usage on GMI Cloud via OpenRouter or the GMI console.

Ollama on Fortune 500 Shift to Open-Source Models

Ollama's jmorgan discussed with Yahoo Finance why Fortune 500 companies are flocking to open-source models for lower costs and greater data control.

Synthesia Launches AI Roleplay Practice Feature

Synthesia's Roleplay Sessions let users practice real conversations with AI avatars that respond, push back, and provide coaching feedback.

GLM-5.2 Vision Version Released, Community Feedback Sought

SGLang confirmed GLM-5.2 now supports visual input. The community is eager to evaluate its performance on UI understanding and OCR tasks.

Greg Brockman Recommends Codex Security Plugin for Cyber Defense

OpenAI's co-founder promoted the Codex security plugin for applying AI models to cyber defense operations.

Mistral CEO: Leaders Moving to Open-Weight Solutions for Deployment Control

Arthur Mensch noted that enterprise leaders are shifting to open-weight solutions to own their AI deployments and intellectual property.

Simon Willison: Loops Were a Short-Term Patch for Model Inadequacy

Simon Willison argued that agentic loops were a temporary fix for models that couldn't reliably sustain work on long problems. Fable, GPT-5.6, and Kimi K3 can now complete such tasks autonomously.

Research & Quick Takes07.24

© 2026 FAV0 · AI Daily