Developer Tools · Product
Cursor Ships Origin, Its Own Code Host
Origin is Cursor's answer to GitHub: fast, deeply integrated with the editor, and synced straight from your existing repositories.
Origin, the code hosting platform from Cursor, is now live. The pitch is speed and deep integration: start by syncing your repositories from GitHub, then keep code, review, and the AI agent that writes it all inside one loop. The launch drew an outsized response within hours, a sign of how closely developers are watching the collision between editors, hosting, and agents. The underlying bet is that hosting is no longer a neutral utility but a first-class part of the coding workflow.
Anthropic
Claude Code Can Now Design
The new /design skill, a research preview, brings Claude Design's artboard workflow into the CLI and desktop, built on artifacts. Run /design to get editable artboards for your UI: pick one, tweak it, then have Claude implement it.
Case Study
ABC Legal Turns Every Employee Into a Builder
A fleet of more than 50 Claude Managed Agents handles specific legal tasks, cutting costs by up to 50 percent. Built-in feedback loops let the agents improve over time, a managed move from ad hoc AI experiments to a governed agent fleet.
The core of open-source AI is that the training recipe itself is the closest analogue to Linux. Model weights are transient.
Nathan Lambert
Infrastructure
NVIDIA to Power Ohio's PORTS-Pike AI Campus
At SB Energy's PORTS-Pike Technology Campus in Southern Ohio, NVIDIA will be the exclusive AI compute infrastructure provider. The framing: AI can transform every industry, but the infrastructure behind it must keep pace.
Enterprise
Replit Adds Governance Tools for the Enterprise
New audit logs cover more than 50 event types, from deployments and identity to secrets and agent activity, with native streaming to SIEM. They join management APIs and workspace settings aimed at easing the burden on IT, procurement, and admins amid AI tool sprawl.
The Agent Loop
xAI
Grok 4.6 Is Smart, Fast, Affordable
Elon Musk sums up the new model in three words: smart, super fast, affordable.
xAI
Grok @Bot
A pointer to the new agentic bot that plugs Grok into everyday workflows.
xAI
Grok Excels at Agentic Tasks
Grok is very good at agentic tasks, Musk writes, pushing the bot as a general agent.
xAI
Create With Grok Imagine
A teaser for image generation inside the Grok family.
Vercel
Origin Repos Deploy to Vercel
Rauchg: host on Cursor Origin, deploy to Vercel, which is itself hosted on Vercel.
Pricing
GPT-5.6 Sol Half Off
50 percent off on AI Gateway through September 18, in both standard and fast mode.
Image Story
The First AI Animated Film Screens at SIGGRAPH
KÖK BÖRÜ blends generative AI with human art, screening at SIGGRAPH 2026 in Los Angeles alongside Pixar, Disney, and NVIDIA. The project is now live and open-sourced, with all prompts and assets available to the community.
Video Models
Seedance 2.5 in 1080p Is Live
Now in 1080p, upscalable to 4K with Luma.
Seedance 2.5 Reaches the Pika API
1080p via the Pika API Club, up to 60 percent cheaper, with sharper textures and cleaner details.
Open-Weights Video, Generated Locally
An otter-on-a-laptop clip made with MiniMax H3 runs entirely on a local machine in about three minutes.
Building Cost-Effective Agents on GPT-5.6
OpenAI shares what startups learn about smarter model selection, reasoning, and tool calling.
Ollama Leads on DeepSeek V4 Flash
Ollama reports the best average performance on DeepSeek V4 Flash; for local use, try the optimized qwen3.8:27b on Apple Silicon or NVIDIA.
SGLang Updates Qwen3.8-27B Recipes
New RTX 5090 and RTX Pro 6000 recipe variants add non-spec, MTP, and DSpark options, with high-throughput and low-latency paths.
Datatrove 0.10.0 Ships
The data-processing library behind FineWeb, FineWeb2, and FinePDFs rolls out a new release.
ExtractBench Scores Grounding Strictly
A field only counts when the value and its citation are correct, at a word-level box with IoU 0.5.
LM Studio's Bionic Gains Agent Skills
Skills are reusable shortcuts capturing a task, specific knowledge, or custom instructions, invoked with the @ syntax.
Nemotron 3.5 Lightning, a Small MoE
NVIDIA expands its Nemotron suite with a 30B-parameter MoE with 3B active, aimed at the smaller-model crowd.
Cumora Puts AI Agents in Your Group Chat
An open-source project that turns agents into chat members with names, personas, memory, and real email.
Self-Verification Beats Claude Fable 5
Scaling self-verification on DeepSeek V4 Flash tops Claude Fable 5 on Terminal-Bench 2.1 while being 11x cheaper.
River API Outperforms Tinker on RL
A blog post tests the River API against Tinker on reinforcement learning runs with identical training code.
Research and Opinion
Opinion
Tool Use Didn't Kill the Need to Scale
Jason Wei reflects: when models first learned to use tools, the tempting narrative was that a 1B-parameter cognitive core plus tools would be enough. The evidence since has complicated that story, keeping scaling at the center.
Agents
Three Ways to Give an AI a Computer
Ethan Mollick maps the field: Codex and Claude Code use your local machine; ChatGPT Work hands the model a one-time online machine that resets; Grok Bot gives each agent a persistent web machine.
Research
Machine Studying: Expertise From a Corpus
A new proposal asks whether an agent can develop domain expertise from a document corpus alone, with no downstream tasks or rewards. The StudyBench benchmark finds frontier models with search still gap on topics emerging after their training cutoff.
Policy
The Claude Watermarking Debate
The key fact in a fast-moving debate: it is possible to watermark LLM-generated text without degrading output quality, which reframes what detection and disclosure could look like.
AI is making better AI. That creates more use cases and more usage, which creates more data and feedback, which helps make better AI.
Michael Dell, via Arav Srinivas
DeepSeek's Price Hike Ripples Out
After winning on price, DeepSeek raised rates: cache up more than 12x, overall 3-6x. Providers are reshuffling token plans.
Is V4 Pro Trained Off-Course?
V4 Flash, several times cheaper, beats Pro, and V4 Pro High beats Pro Max, though Pro still leads on prose style.
Functionally Flash
Very few evals show Pro significantly exceeding Flash, raising questions about whether the V4 architecture scales.
Dense Models Drag on Local Hardware
Dense models are slow locally, but the progress is a sign of things to come soon-ish.
Physical AI Evals Become a Category
VLA evaluation is moving past simple suites like LIBERO into a fuller physical AI benchmark category.
Models Learning From Their Own Data
It is now routine for models to generate their own training data and learn from it, something unclear even a year ago.
RL for Large MoEs, Zero Mismatch
Researchers can now RL large MoEs with zero train-infer mismatch, and it improves performance.
We Underutilize Our GPUs
A reminder that the industry is still far from squeezing its existing compute.
Confucius4-TTS Upgraded
The Confucius4-TTS paper is on arXiv, and the open-source model just received a major upgrade.