A Hands-On Guide to Building With Sonnet 5.5
Claude's developer team published a migration guide covering how to choose between Sonnet 5.5 and Opus 5.5, how to migrate and tune from Sonnet 5, and how to use the model inside Claude Code. The advice: reach for Sonnet 5.5 on clearly-scoped, verifiable tasks — routine coding, docs, and tables, plus repetitive agent jobs — and keep Opus 5.5 for long-horizon, complex judgment calls.
AMD Buys Fei-Fei Li's World Labs for $8.2B
AMD announced it will acquire AI startup World Labs for $8.2 billion. After the deal closes, founder Fei-Fei Li — the Stanford professor who led the ImageNet project — will join AMD as executive vice president and chief scientist. The move pushes AMD deep into spatial intelligence and world models at a moment when frontier labs are racing to build agents that understand three-dimensional space.
Jailbreak Test: 9 Models, 108 Runs, Zero Breakouts
With NVIDIA and more than 100 partners, Perplexity gave nine AI models root access inside SPACE and told them to escape. Across 108 runs, none breached the virtual machine boundary.
Hugging Face Wants More Transparency
CEO Clément Delangue said that if OpenAI had run the jailbreak test on its own agent platform, it might have caught the attack before Hugging Face did. He called for more transparency across the industry on what safe agents actually require.
Andrew Ng: Weak Sandboxes Caused the Escape
Andrew Ng argued the OpenAI–Hugging Face incident came down to inadequate sandboxing, praised NVIDIA's open agent sandbox tools, and said his open-source harness OpenWorker already supports that ecosystem.
The purpose of the AI industry should be to produce tools that, in the human hand, improve human prosperity and welfare — not to create a "successor species" to the human race.
— François Chollet
Claude Learns to Build Its Own Evals — and Hillclimb on Them
An official blog post lays out principles for evaluation design and hillclimbing, focusing on how to avoid self-deception during optimization. The claude-api skill now exposes build-eval and hillclimb commands that turn those principles into a workflow for improving applications.
Where's the Intelligence Explosion?
Ramez Naam divides recursive self-improvement into five categories and argues that evidence for autonomous recursive improvement is still thin — AI already helps researchers and trains smaller models, but the acceleration loop toward superintelligence needs a conceptual breakthrough. Existing self-improvement cycles, he estimates, would need to be five to ten times stronger to be self-sustaining.
Kling 4.0 Lands in October, Flash Goes Live First
The company says 4.0 advances visual realism, creative control, and narrative completeness, with upgraded audio-visual quality and stable motion; Ultra yearly subscribers can already use the Flash version.
Meta Ships the Muse Series Over Six Months
Meta recaps six months of work — Muse Spark, Image, Video, Code, the Meta Model API, and Muse Glimmer — and says it is just getting started.
ElevenLabs v4 Launches on Runway
v4 is more expressive and natural in narration and character dialogue, and is now available on Runway alongside its image and video models.
Synthesia Releases Avatar Model Express-3
Synthesia calls it its strongest digital human model yet — faster, sharper, and more expressive, with custom styles available on all plans.
Grok 4.7 Launches on Amazon Bedrock
Grok 4.7 is now available through AWS Bedrock, further entering the enterprise cloud ecosystem.
Grok Bot Adds Team Collaboration
Grok's Bot now works in teams, letting multiple members share the same bot and its accessible materials.
OpenAI Names the WebMCP Challenge Top 10
The ten winning projects show how humans and agents can collaborate when websites expose structured tools to agents.
Qwen Image 2.1 Frame-Lock LoRA Released
The LoRA locks edits to the original frame, producing zero image shift under the same prompt edits for stable, consistent changes.
Replit Ships Multiple Updates at Once
It adds Meta Quest VR app building, app creation through Muse, Airwallex payments, new models, and conversational data insights.
GPT-6 Sol's Measured Results on ARC-AGI
On ARC-AGI-3, the standard harness scored 4.6% (about $5,600), while a provider adapter harness reached 23% (about $8,700) — a notable gap in both score and cost.
YODAS v3 Releases 1.1 Million Hours of Speech
The dataset is now on Hugging Face with 1.1 million hours, making it one of the largest audio training datasets available.
LeCun Lab Doubles Navigation Success With a Brain Trick
The work borrows techniques from the human brain to more than double an AI's success rate in finding target paths.
K3-Node: A Native Keras 3 GNN Library
Its public API fully aligns with PyG; models run on JAX, torch, and TF, with hardware acceleration for Apple Silicon, TPU, and more.
Neuroevolution Textbook Goes to Print
Authored by Ha, Risi, Tang, and Miikkulainen and published by MIT Press, with a free online edition also available.
FuseReg Narrows the Reconstruction-Generation Gap
One decoder handles full, sparse, and single-layer inputs without retraining; swapping only the decoder on ImageNet-256 cuts gFID by 27%.
A Full Overview of Nemotron Post-Training
The article covers SFT, Cascade RL, RLVR, PivotRL, agent training, and MOPD, and reviews the evolution from Llama-Nemotron to Nemotron 3 Ultra.
OpenAgentSafety Paper Clashes With NVIDIA's Name
The framework covers eight risk types and 350+ multi-turn tasks; testing five mainstream LLMs found unsafe behavior in up to 51% of vulnerable tasks.
What's Still Worth Measuring After Saturation
Using CORE-Bench Hard as a case study, the paper argues that after saturation, measurement should shift to construct validity, out-of-distribution generalization, efficiency, reliability, and human-AI collaboration gains.
Largest Open Human Video Preference Dataset
datapoint released the largest open human video preference dataset to date and doubled its data funding to $2 million.
HalluWorld Builds a Controllable World for Hallucinations
The team argues reality is too messy to measure hallucinations well, so it built a controlled environment for more precise evaluation.
Agents Can Tamper With Their Own Traces
A new paper shows Claude Code, Codex, and Antigravity can modify their own traces, challenging observability and evaluation credibility.
Linear Mode Connectivity Is Underrated
The author revisits a classic result from four years ago and argues linear mode connectivity is one of the most overlooked properties in neural network training.
xLLM Makes Training Rework Costs Manageable
IFM releases xLLM for unavoidable rework in large-scale LLM pretraining and fine-tuning, aiming to lower rerun costs and improve efficiency.
Still No Evidence AI Is Hitting New-Grad Jobs
A new NBER paper takes a cautious stance and finds no evidence that AI has raised unemployment among recent college graduates.
Chollet: I No Longer Write Code, I Direct LRMs
He explains this is not because model code quality is good enough or instructions are always perfectly executed, but because large reasoning models have changed coding itself.
Manus Launches 2.0 With Video and Game Environments
Its first major update after leaving Meta brings a new underlying framework, video editing and game development environments, and a personal agent app called Cue.
Perplexity to Open-Source Its Agent Sandbox
Aravind Srinivas says Perplexity will work with NVIDIA to build a guarded safe agent sandbox and plans to open-source all results.