Claude Opus 5.5 matches Fable 5.1 — for less
Anthropic released Claude Opus 5.5, the first model in the Claude 5.5 family, and claims it performs at the level of Fable 5.1 on most work. The model is about 30 percent faster and 40 percent cheaper per task than Opus 5, with input and output priced at $4 and $20 per million tokens. Claude Code users also get a 20 percent increase to the five-hour session limit today, along with reset quotas for Pro, Max and Team accounts.
Especially compared by per-task pricing, which is the metric that should matter, I don't think there is anything competitive anywhere in the market.
GPT-6 Sol and Luna hit the API
Both models go live today at prices 50 percent lower than GPT-5.6, aimed at agents that need to run fast and cheap in production.
Xiaomi open-sources MiMo-V2.6
Pro and Flash are natively multimodal. Pro scores 46 on the Artificial Analysis intelligence index, beating Kimi K3 and Qwen3.8 Max to tie Grok 4.7 as the strongest open model.
Kimi K3 lands on Amazon Bedrock
Moonshot's strongest open-weight model arrives on Bedrock with a one-million-token context, vision, and explicit prompt caching for coding and agent workflows.
GPT-6 prompt caching upgraded
Higher cache-hit rates are now the default in the GPT-6 API, with cached-input discounts of up to 90 percent so agents run faster and cost less.
Claude Opus 5.5 arrives in Cursor
The new top model on CursorBench at 57.8 percent, and 40 percent cheaper per task than Opus 5.
GPT-6 Sol now available
Sol becomes the default Light option in Computer's effort selector.
Computer adds Opus 5.5
Opus 5.5 scores 0.610 on WANDR at $4.13 per task, 67.6 percent cheaper than Fable 5.1.
A browser extension, not a WebBridge
Chat, navigate and fill forms from the sidebar, and record repeated steps as reusable skills.
Grok Bot arrives in Tesla
Musk says Grok Bot is now live in Tesla vehicles.
Step Code v0.1.0, MIT-licensed
A CLI coding agent covering read, edit, test and ship, scoring 80.9 percent on one benchmark.
DeepSeek Elastic Compute
A single production-scale unit of DSec spans about 160 nodes and serves roughly three million sandboxes per day. In production it supports over 380,000 concurrent sandboxes and sustains more than 5,000 sandboxes per second.
bitsandbytes2 hits 1.5 to 2.0 bits
Tim Dettmers' runtime dynamic compression framework reaches 1.5 to 2.0-bit compression at high quality, integrated into the bitsandbytes2 library now entering a private beta.
RoboDawn: VLM intelligence for robots
A frozen VLM controls robots in closed loop; one-shot success on RoboTwin rises to 73.6 percent versus 46.0 for π0.5.
RRSI self-improves agent harnesses
Recursive self-improvement of prompts, control flow, tools and memory under a frozen base model, with constraints to fight overfitting.
WorldCrafter video world model
Implicit 3D-aware memory lets viewpoints govern compression, enabling streaming scene exploration from a single image or text.
CodeMidas scales RL from code
From source alone it generates 5,545 tasks across 3,185 libraries and 23 languages; GRPO-trained MiMo-V2.5 gains up to 17 percent.
Perplexity teaches Computer from its mistakes
Hint-guided self-distillation cuts tool-call failures by 21.2 percent in live A/B tests.
vLLM drives RL on 1,000+ TPUs
peano_ai runs full-parameter RL on MiMo-V2.6 310B; vLLM rollouts move all 310B parameters across the ICI fabric in under two seconds.
Step 5 Preview resets the frontier
Artificial Analysis intelligence index of 44 at $0.71 per task, with more releases promised October 15.
Conformal prediction, gently introduced
A distribution-free uncertainty method that yields statistically rigorous confidence sets for any model, with code and extended tasks.
onPanda: token-level correction
StepFun open-sources its annotation tool; correcting the first bad token cuts median annotation time 52 percent versus manual editing.
v0.5.20 adds Intel XPU
HiCache L3 warms cold starts
LiteParse 25% faster
funes local memory
DeepSeek billing fixed
Qwen-Image-2.1 on OpenVINO
WebCraftBench live apps
Knowledge work will industrialize like manual labor did
Ethan Mollick argues the industrialization of knowledge work will be as disruptive as the industrialization of physical labor, with craft fields pressured to produce volumes of standardized output using less of the craft that made the work meaningful. The coming shift, he warns, is as much about meaning as it is about productivity.
Andrew Ng calls out AI panic hype
Ng says the recent escalation in AI danger panic was not driven by a change in the technology itself, but by hype — possibly an orchestrated PR campaign — and calls it a step backward for the field.
OpenAI opens up to independent evaluators
The lab pledges deep third-party access across training, evaluation and deployment so assessors can challenge assumptions and surface risks.
Runway launches DIFFUSE talent platform
A marketplace for agencies, brands and studios to find and hire AI-native creative talent.
Next.js benchmark ties at 97%
Opus 5.5, GPT-6 Sol and Fable 5.1 all score 97; Grok 4.7 gets 94 at a fraction of the cost.
A same-day price war
A review of Opus 5.5 and GPT-6 Sol/Luna, comparing reasoning levels with a pelican SVG.
Altar-1 security model
The first open-weight defensive security model, built to deploy and own.
Swap Anything app
Replace cast, wardrobe, setting or props while preserving timing and motion.
Relight Media
Rebuild light and shadow on any photo or video while keeping subject and framing.
AI shortens R&D to 2–3 days
A reported 10× productivity gain; Xiaomi needs industrial AGI for its hardware stack.
Recursive self-improvement
Qwen explores models that find weaknesses, design experiments and keep training.
Sol and Luna, half the price
Altman frames the pair as a major leap on intelligence, coding and computer use.
A $50–60B AI factory
Jensen Huang warns gigawatt-scale factories can't afford overly specialized architecture.
MiMo-V2.6 tops the board
Simple GQA plus a 128-token sliding window ranks first on weighted-average open-weight.