Model Release · Mistral AI
Mistral launches “Le Chonk,” a 1T-parameter natively multimodal flagship
Code-named Le Chonk, Mistral Large 4 activates 49 billion parameters per token and is billed as the strongest open-weight model to come out of Europe or the US.
Paris-based Mistral AI unveiled Mistral Large 4, its new flagship, and gave it a nickname that has already stuck with the community: Le Chonk. The model packs one trillion total parameters with 49 billion active per token, and the company positions it as the best open-weight model from the United States or Europe on aggregated benchmarks. It is natively multimodal and, Mistral says, reaches state-of-the-art results on critical workloads including cybersecurity, manufacturing, and finance. The open weights are expected to land by the end of October — a release that would crown the most consequential Western open-model launch of the season.
Research · OpenAI
OpenAI publishes new math results from its internal frontier model
The lab turned to the Institute for Advanced Study for review before publishing.
OpenAI released a broad set of mathematical results produced by an internal frontier model. In an unusual move for a commercial lab, the company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and drew on its advice before publication — a sign that frontier labs are reaching beyond benchmark tables toward substantive, peer-shaped scientific contributions.
Product · OpenAI Devs
Decisions API opens to every developer in public beta
The Decisions API lets an app choose the right model, tool, or action in near real time. OpenAI says decisions run up to 10x faster than GPT-6 Luna through the Responses API.
Model Release · Google DeepMind
EmbeddingGemma 2 unifies text, code, image, audio, and video on-device
Google’s first natively multimodal open embedding model maps four modalities into one shared vector space, targeting efficient on-device deployment. Day-0 support landed the same day across llama.cpp, vLLM, and SGLang.
Security · Anthropic
Anthropic widens its cyber verification program
Verified security practitioners gain broader access to Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5, with safeguards tuned for defensive work. The expansion is the clearest signal yet that defensive cybersecurity is becoming a first-class workload for frontier labs.
“The cyber-risk discourse is broken — critics ignore that closed models cause most existing attacks.”
Steer desktop agents from your phone
The Cursor iOS app now lets you check progress, reply, or start new coding tasks while away from your machine.
API rate tiers shrink from five to three
Paid tiers simplify to Build, Launch, and Grow; the top tier now unlocks at $500 in cumulative spend, down from $1,000.
Claude Code cloud sessions run in parallel and resume
Each task runs on a fresh VM, so sessions keep going when the laptop closes; a field guide covers seven workflows.
Mistral touts Large 4 for cybersecurity
It scores 82% on CyberGym-E2E-AA, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (78%).
Replit connects context across every project
A new chat can find an existing project, add a feature, or reuse one project as a reference for another.
Perplexity open-sources decision model v1.1
pplx-decider-v1.1-27b is an updated open-weight multimodal decision model that Perplexity says scores highest on its evaluations.
GLM-5.3 lands on Amazon Bedrock
A 744B-parameter MoE flagship for software engineering and agentic work, with ~40B active per token, 1M-token context, and up to 128K output.
Can models compose now? Opus 5.5 passes the test
Simon Willison asked Opus 5.5 to write computer-game music; the result beat expectations, echoing the models’ recent 3D-graphics leap.
H-JEPA learns hierarchical world models for planning
An end-to-end world model that plans from coarse goals to fine actions, lifting Visual AntMaze success from 18% to 73% with less compute.
SGLang adds Kandinsky 6.0 video generation
Video and synchronized audio from text or image, with speech, sound, and lip-sync from one model in Lite (3B) and Pro (29B).
IP made sense when alternatives were costly — AI changes that
François Fleuret argues intellectual property rules make sense when copying is expensive, but AI drives that cost toward zero.
Decision API pricing cut in half
Perplexity’s Decision API drops to 2 cents per million input tokens.
Voice input arrives
Talk to Codex and give your keyboard a break.
No model to rule them all
Ideation and final frames need different models; the job is routing each task to the right one.
Decider is the best decision model
Aravind Srinivas makes the claim plainly, days after the v1.1 release.
RL shows no sign of saturation
Arthur Mensch on Large 4: trained and served on their own compute, with no saturation in reinforcement learning.
FLUX 3 billed as the strongest world model
Black Forest Labs claims the top world model, with an image to match.