Auto becomes the default permission mode
Anthropic rolls out Auto mode as the default in Claude Code for Pro, Max, and Team users. If you have already set a default, Claude will ask before changing anything. Shift+Tab still switches modes on the fly, and a defaultMode setting pins one permanently.
Gemini 3.7 Flash reaches Pro and Ultra
Gemini 3.7 Flash is now available to all Pro and Ultra users in Gemini chat. The update delivers improved reasoning, framed by GeminiApp as the fastest route to the new Flash tier inside the chat app.

206 tok/s on a single RTX 5090
Day-0 SGLang support decodes Qwen3.8-27B at 206.1 tok/s on one RTX 5090 and 38.28 tok/s on DGX Spark — a new bar for what a small model can do on a single card.

vLLM lands Day-0 for Qwen3.8-27B
A 27B hybrid-attention model with 262K native context expandable to 1M, ready for single-GPU deployment in vLLM the day it ships.
By next year, using a computer will be optional. Work will radically change.
DeepSeek-V4-Pro ships MIT-licensed
A big jump in agentic capability over the preview, with the same architecture so your config carries over untouched. vLLM has run this path since 0.25.0.
dots3-note preview: a 280B agentic model
A 280B-total / 16B-active multimodal model with 512K context, built for real-world long-horizon agent tasks — the model behind RedNote's IMO 2026 gold.
Ollama adds the DeepSeek Harness
Run it entirely in your own environment, with web search pre-installed and a trajectory view that shows what is happening in the background.
Anthropic explains its AI watermarking
Watermarking is implemented to comply with the EU AI Act, and other major model developers signed the same Code of Practice and are following suit.
The second Risk Report is out
A routine disclosure under the Responsible Scaling Policy, sharing detailed information on system risks and how prepared the company is to address them.
Small models still dominate local inference
The State of Open Models shows frontier models getting larger while small models dominate real-world use. Qwen leads local inference, followed by Gemma, and AI agents are a major Hub force.
Build flexibility into your AI stack
Ethan Mollick warns that extrapolations about AI's corporate impact are too confident, given how unstable today's prices, adoption, and capabilities remain.
Vision models understand the structure of the world projected onto an image far more than LLMs understand the world projected into words — we should tap into the former to do better abstract thinking with the latter.
Four frontier audio models, 20× cheaper
Soundtrack, Music, SFX, and Speech cover the full spectrum of generative sound, priced below every audio model on the market.
Seedance 2.5 makes 1080p commercials
Full 30-second commercials from a single prompt, with free 1080p generations for new users at zero credit cost.
The Search SDK goes agentic
The SDK behind Perplexity Computer is now available inside any agentic harness, for wide and deep research.
A coding harness nearly solves ARC-AGI-3
Adding a coding harness nearly solves ARC-AGI-3 — coding generalizes LLMs, exactly as predicted.
Testing GLM 5.3: precise and concise
Prefill around 1 ktok/s and output near 60 tok/s. In his harness, full thinking traces already put GLM 5.2 above Fable and Claude Code — GLM 5.3 is another level.
The ARC-3 public set is a demonstration set
Public ARC 3 games are neither eval nor training data; scores on them do not predict scores on the hidden set.

