Policy · Provenance
OpenAI extends provenance to text for the EU
Content verification moves beyond images and audio as regulators tighten the AI Act.
OpenAI is expanding its content-provenance tooling to include text, a step it says responds to EU regulatory requirements under the AI Act. The company is candid about the limits: current text-watermarking technology remains unreliable under editing and paraphrasing, and it is framing the rollout as a first step rather than a solved problem. Existing tools already let users check whether an image or audio file was created with OpenAI models.

Enterprise · Platform
Cohere launches North 2 with 15-plus new features
Cohere frames North 2 as the end of enterprise compromise: a platform that is scalable, sovereign, and secure enough for agentic AI. The release bundles more than fifteen new features in what the company calls its biggest upgrade yet, aimed squarely at teams that need control without sacrificing model quality.
“The clearest takeaway should be that the Chinese are very, very good at building LLMs.”
Industry · Competition
Meta and Microsoft tighten internal use of Claude
The Information reports that both Meta and Microsoft are steering employees away from Anthropic's Claude and toward their own models and tools. Inside Meta, the number of employees using Claude Code fell from roughly 60,000 at the start of the year to about 30,000 recently, a drop partly attributed to spring layoffs but also to an explicit push toward in-house alternatives.
Security-One 27B for security decisions
An open-weight 27B model that reads documents and issues tool calls, tuned for security decision-making.
Cursor SDK steers agents at runtime
A new run.steer() injects messages into the next turn; a busy subagent drops to the background and keeps working.
First audio to 23ms on RTX 5090
fractalyze.io optimizes Qwen3-Omni on vLLM-Omni, cutting time-to-first-audio from 213ms to 23ms with AWQ-4bit at batch one.
llama.cpp v0.6.0 lands
The new release adds Clef support across text and vision, plus high-quality support for Qwen3.8-Flash-Next and others.
Command Code ships open decision model Agr
Agr (31B) and Agr-flash (360M) arrive, scoring 58.15 on Decision Index 0.2.
Reflection opens a Hugging Face org
The official organization page says Reflection builds open models so anyone can control their own intelligence.
Paper
PixelUMM: encoder-free unified multimodal
PixelUMM understands and generates images and video directly in pixel space, dropping VAE and ViT for a Mixture-of-Transformers backbone.
Paper
RMD cuts long-horizon video drift
Rollout-Marginal Distillation scores each clip independently against a clip teacher, so quality fixes no longer chase artifacts from imperfect context.
Paper
RSR expands 759 wins into 11,094 trajectories
Recursive self-rewriting uses a base Qwen model as planner, critic, and executor, growing a handful of solved tasks into a large fine-tuning set.
Paper
OctLLM speaks in octrees
Octree occupancy tokens represent geometry explicitly, with sparse S-Octrees balancing shape fidelity against sequence length without breaking the language pathway.
Paper
Full-bandwidth transformer gains vertical feedback
A new revision fuses previous-layer hidden states with sampled tokens, preserving standard architecture and KV cache while improving several 1B evaluations.
Experiment
Qwen3.8 27B adds long numbers locally
Simon Willison reruns the GPT-4o long-number addition test against local Qwen3.8 27B, in both reasoning and non-reasoning modes.
Engineering
Symmetric memory speeds up NCCL
Recently added to NCCL and PyTorch, symmetric memory delivers faster small-to-medium payload comms in a few lines of code.
Pricing
OpenRouter warps inference economics
For GLM 5.3, the highest-volume provider charges $0.08 per million input tokens but $5.00 per million output — a brutal prefill-decode inversion.
Industry
Claude delivers 5x the tokens for $200
SemiAnalysis tests nine subscription plans and finds Claude's value on its flagship model runs about five times that of OpenAI at the same price.
Market
The rise of the AI whale
The top 1% of consumers spend more than $900 a month on AI products on personal credit cards, outspending the bottom 50% combined.
Product
Codex CLI learns to take voice
OpenAI demos launching tasks by voice, managing agents across projects, and exploring directions in isolated worktrees.
Product
Hedra opens a shared team canvas
Every teammate works beside their own Hedra agent on mood boards, prompts, and references — ending the era of final_final_v3.png.
Product
Higgsfield's ROAS-driven ad engine
Ads Studio researches a brand, tracks trending formats, and drafts static ads from industry best practices.
Model Release
Recraft V4.1 chases the spontaneous shot
Natural photography, stronger composition, and authentic everyday moments turn simple prompts into images that feel alive.
Commentary
Kai-Fu Lee: AI needs a new company
After talking with 100-plus CEOs, Lee argues the first mistake is assuming we have seen this before; AI demands a different kind of organization.
Tooling
lcu frees Codex Computer Use
The project extracts Computer Use so Claude Code, Codex CLI, and Pi can also drive a screen, mouse, and keyboard.
Hardware
Is CUDA really NVIDIA's moat?
A shared view argues CUDA developers can move to Ascend C without much pain, letting latecomers catch up as hardware and software iterate together.
Analysis
Beam, without the taint of Claude
The excitement is that Beam appears completely undistilled, offering a rare look at competitiveness built from clean data.

