DeepSeek Ships V4.1-Flash, a Faster, Sharper Model Family
The smallest model in a brand-new architecture ships with native vision, higher throughput, and a clean path to scale.

DeepSeek has released V4.1-Flash, the smallest model in a brand-new architecture family and the first with native visual understanding. The company says it is more capable, reasons faster, and pushes more tokens through the same hardware, all while staying cheap to run.
The architecture is built to scale: the same design that powers the Flash tier is meant to grow into larger models. Early reactions from researchers describe the release as a rewrite of the Transformer playbook, dropping HCA, generalizing CSA, and rethinking how context is compressed and retrieved.
GPT-Live-1 Hits the API, Bringing Real-Time Voice to Apps
Voice agents that listen while they speak, paired with the model and harness you choose.

OpenAI has opened GPT-Live-1 to developers through the API, porting ChatGPT's natural back-and-forth speech into applications. Agents can listen and speak at the same time, and teams can pair the model with whatever harness fits their stack.
OpenAI Launches a Financial-Services ChatGPT

A tailored ChatGPT Work experience combines built-in financial data with GPT-6 Astra reasoning, letting teams research, build financial models, and produce custom client materials.
OpenAI Opens the Agents API in Public Beta

Developers can build and run cloud agents on the Codex harness, fully managed by OpenAI. Orchestration, long-running sessions, and context management are handled for you, so you can focus on what makes the agent unique.
With its new architecture, this release is just as big as R1 and will redefine all future models.
— Tim Dettmers, on DeepSeek V4.1

Tencent Hunyuan Releases the Open Audio Model AuK
AuK is an open-source foundation model for unified speech generation and editing. One interface takes natural-language instructions plus reference audio, covering zero-shot TTS, instruction-controlled generation, content editing, and voice conversion.

Cohere Ships North Small Translate, an Open Translation Model
Nine years after the Transformer was proposed to improve Google Translate, Cohere returns to the task with a leading open machine-translation model, with weights published for download.
Cursor Launches Projects, a Persistent Coordinator

Instead of a fresh chat per task, one long-lived thread with a coordinator agent manages work through subagents and improves over time.
A Data Agent Joins ChatGPT Work
Teams can now put company data to work inside ChatGPT, connecting internal sources for analysis and processing in a single workspace.
GPT-Live-1 Reports Voice Benchmarks

OpenAI shares results on task completion, back-and-forth turn-taking, response speed, and tool use for production voice agents.
SGLang Ships Day-Zero Inference for V4.1-Flash
The 552B backbone pairs shared compressed KV, a two-stage sparse indexer, and 196B Engram memory.
vLLM Serves V4.1-Flash on Day Zero
Verified on NVIDIA and AMD GPUs: 552B MoE, native vision, 1M context, 8B active on read, 16B on write.
Ollama Rolls Out V4.1-Flash
Max and Team accounts first, with capacity expanding to all subscribers.
Chollet Defines "Symbolic Learning"
Machine learning on a symbolic substrate — learned functions look like code, not curves.
Raschka: V4.1 Deserved a V5
The encoder-decoder overhaul is big enough to earn a new major version, he says.
YOCO: Cache the KV Only Once
A self-decoder plus cross-decoder reuses one global KV, cutting memory while keeping global attention.
Sakana AI, Sumitomo and SCSK Team Up
Sakana's tech, SCSK's systems, and Sumitomo's 100,000-company network push domestic AI into Japanese firms.
Schulman: User Data Adds Little
Frontier math gains come from pretraining and RLVR; user data mainly surfaces failure modes.
FLUX 3 Video Edit Is Live on fal
One prompt swaps a character, rebuilds the background, or restyles a shot.
How One Resignation Set AI Fear Ablaze
Lambert disentangles which risk claims hold up and which are hyperbole.

Hooker on the Slow Death of Scaling
Returns from piling on compute and parameters are flattening.
Mollick: Math Previews the Jagged Frontier
What hits mathematicians now reaches other professions next.
Jan Leike Wants a Brake on the Frontier
Institutions should slow the scaling race to buy time for safety and alignment.
Lambert Proposes "Lossy Self-Improvement"
Research acceleration is real but friction-bound; progress may be linear, not explosive.

Astra's CoT-Free Compute Jump
Astra does 1.75x the steps of its rivals without chain-of-thought — a concerning trend.

mmap Can Be 4,937% Slower
Beware mmap on network filesystems; mitigation in the ML-engineering notes.
A 40-Layer Teardown of V4.1-Flash
Twenty causal encoders feed a decoder whose global KV comes from encoder final states.
V4.1 Rewrites the Transformer
HCA ditched, CSA generalized, sparse-attention warmup dropped.
Claude Agents Get a Session Viewer
The ant CLI attaches a terminal to a running session, or opens a local web UI.
Claude Code Pops Panes Out
Drag a diff or terminal to a second screen and dock it back anytime.
A 100-Agent Study on Group Drift
How a small minority can shift a crowd's answers on math problems.
Runway on Real-Time Video Research
After Solaris and GWM Worlds 2, instant generation is where video is headed.
Replit Adds Scheduled Routines
Hourly, daily, or weekly runs start with deterministic code, calling the Agent only when needed.
Higgsfield Effects 2.0 in ChatGPT
Type @Higgsfield /effects, upload a photo, get a viral video — free for now.