Kimi K3 Launches on Databricks
Moonshot AI's flagship model arrives on Databricks, bringing enterprise-grade reasoning to the platform.
Moonshot AI's Kimi K3 is now available on Databricks for enterprise use. The integration brings one of China's leading large language models to the Databricks ecosystem, enabling enterprise customers to run advanced reasoning workloads directly within their data infrastructure. Kimi K3 has been recognized for its strong performance on coding, long-context reasoning, and multilingual tasks.
Hy3D WorldClaw: Generate 3D Worlds from Text Prompts
Tencent Hunyuan unveils an agentic workflow that builds explorable, game-ready 3D open worlds from text.
Tencent Hunyuan unveiled Hy3D WorldClaw, which generates large-scale, freely explorable 3D open worlds from text prompts, built entirely from editable game-grade assets. Unlike video or Gaussian Splatting approaches, WorldClaw produces structured 3D environments that users can freely navigate and modify, marking a significant step toward production-ready AI-driven worldbuilding for games and simulations.
ChatGPT Desktop App Arrives on Linux in Preview
OpenAI brings ChatGPT, ChatGPT Work, and Codex to the Linux desktop, meeting developers where they already work.
OpenAI released a preview of the ChatGPT desktop app for Linux, supporting ChatGPT, ChatGPT Work, and Codex. The app integrates with existing projects and browser workflows on supported Linux systems, closing a long-standing gap for developers who prefer the Linux ecosystem. The announcement garnered over 8,100 likes and marks one of OpenAI's most anticipated product expansions this year.
Mistral Sets a Roadmap for European AI Autonomy
In-region inference, open models, and new European compute infrastructure for sovereign AI.
Mistral announced it is integrating inference infrastructure, open models, and long-term commitments to help Europe take control of its AI future, setting a global development roadmap that emphasizes in-region inference and new compute for sovereign AI.
"We were promised flying cars and all we got is infinite superintelligence for everyone."@rauchg, CEO Vercel
ChatGPT Desktop App Supports Cross-Agent Import Sync
Import projects, chats, skills, and plugins from other agents with automatic sync updates.
OpenAI now allows users to import projects, chats, skills, and plugins from other agents, with automatic sync updates available in the desktop app. Users can review import history and opt in to automatic updates in Settings, enabling seamless workflows across multiple AI coding assistants.
Codex Lands on Linux with ChatGPT Desktop
Build where your code already lives: Codex is coming to the ChatGPT Linux desktop app.
OpenAI announced that Codex will arrive on the Linux version of the ChatGPT desktop app, letting developers build directly where their code lives. The integration keeps projects, development workflows, and supported browser tools together in one native application.
SGLang Ships Day-0 Support for Nemotron 3.5 Lightning
SGLang launched Day 0 support for NVIDIA Nemotron 3.5 Lightning, a 30B hybrid Mamba-Transformer MoE with just 3B active parameters, distilled from Nemotron 3 Ultra for always-on agents. The release includes three built-in speculators — MTP, DFlash, DSpark — and native support for coding agents, tool use, and agent harnesses.
Nemotron 3.5 Lightning Now Available on Ollama
NVIDIA's 30B model for always-on agents is now on Ollama, running entirely local. Launch commands for Claude Code, Hermes Agent, and OpenClaw are included out of the box.
Nemotron 3.5 Lightning on LM Studio
The 30B MoE model with 3B active parameters runs fast on LM Studio, supporting up to 1M context window. Distilled from Nemotron 3 Ultra for high-volume agentic use cases with tool use support.
"A great American open weights MoE model that can run efficiently on your laptop or local hardware like the DGX Spark. You can use the larger Nemotron Ultra on Perplexity."@AravSrinivas, CEO Perplexity
Gemini App Surpasses 1 Billion Monthly Users
Google's fastest-growing product ever and the company's 14th to hit the billion-user mark.
CEO Sundar Pichai announced that over one billion people now use the Gemini app every month, making it Google's fastest-growing product ever and its 14th to cross the billion-user threshold. Pichai credited Josh Woodward and the entire Gemini team for the achievement.
US Open AI Model Development Is Back
Graham Neubig of CMU notes that back-to-back 30B releases from Meta and NVIDIA signal a resurgence in open US AI development. "This is great for local AI, given that most other labs have been building bigger and bigger models that are not easy to run locally."
ZCode Hits 1 Million Users
ZCode reached 1 million users and reset usage limits for all GLM Coding Plan users. An update improves long-horizon capabilities in real engineering workflows with 98% cache efficiency.
Seedance 2.5 Goes Live on Runway
Runway launched Seedance 2.5 with 50 unique character references and clips up to 30 seconds synced to music, supporting full cast and track generation.
Luma Scenes: Refine Before You Render
Luma Labs unveiled Luma Scenes, powered by Uni-1. Users refine every scene and approve everything before rendering, eliminating wasted re-generations when one wrong scene ruins a production.
Wan-Animate-2: Major Open-Source Animation Upgrade
Alibaba Wan released Wan-Animate-2 with high-fidelity character animation, multi-character support, text-controlled camera viewpoint, and real-time streaming generation. Weights and docs are available.
Wan CLI Ships for Programmatic AI Video
Alibaba Wan released a CLI for calling Wan from agents to generate images and videos programmatically. All CLI credits sync with the Wan video account.
ExtractBench: The Most Comprehensive Document Extraction Benchmark
LlamaIndex tested 14 systems — frontier VLMs, coding agents, extraction APIs — on 370 enterprise documents across 4,869 pages and 67 document types.
Meta Muse Glimmer: First New Open-Weight LLM Since Llama Era
Meta released Muse Glimmer, a 30B multimodal reasoning model with a Gemma-like architecture. Weights are available on Hugging Face with BF16 weights, GGUF k-quants, ExecuTorch builds, and DFlash draft models.
Nathan Lambert: "We Don't Need Policy on Distillation, We Need Products to Work"
"When I was in China it was insinuated that every lab did this. I'm glad there's public research on it and am still shocked the frontier labs haven't patched this stuff. We don't need policy action on distillation, we just need the products to work as intended." The thread resonated widely with 677 likes and 66K views.
vLLM: 4x Throughput for Nemotron Lightning
Run Nemotron 3.5 Lightning on vLLM with up to 4× higher throughput and 30% faster task completion. OpenAI-compatible API for easy agent stack integration.
SGLang Introduces Unified Radix Cache
The cache class matrix is dead — Unified Radix Cache handles full attention, sliding window, and recurrent state caching for hybrid models in agentic serving.
SGLang Powers Nemotron NVFP4 Local Deployment
SGLang with DSpark enables Nemotron 3.5 Lightning to run with NVFP4 and speculative decoding on DGX Spark, RTX 5090, and RTX 6000 PRO.
Slime Open-Sources GLM-5.2 Training Alignment Path
The RL framework behind GLM series training achieves deterministic train-rollout alignment with logprob MAE of 1.9e-7 at 4096 tokens.
Ideogram Ad Resizer Exports Campaign-Ready Sets
Upload one design, choose placements, and export for Instagram, Facebook, YouTube, connected TV, display, and print at extreme aspect ratios.
Recraft V4.1 Maintains Style Across Diverse Subjects
One frost-crystal style holds material logic through a crown, jewels, a knight, and a jet — demonstrating strong style consistency.
Higgsfield Layers: Non-Destructive Layout Edits
Automatic layer decomposition for poster edits with undo at every step. Layouts stay flexible after generation.
Higgsfield Layers: Turn One Poster Into a Campaign
One flat poster becomes a full multi-variant campaign with flexible, editable layout after generation.
Grok Bot Enters Early Beta
Grok Bots are AI teammates that sign in to your tools and do real work for you, now in early access.
Elon Musk: Grok 4.6 Ships Later This Week
Grok Bot beta will widen after basic issues are fixed and Grok 4.6 is released later this week.
Grok Imagine Powers Steampunk Movie Generation
Elon Musk demos Grok Imagine, showing users can generate their own steampunk movies with the tool.
Ollama and NVIDIA Collaborate on Open Models
"Open models have no boundaries — let's continue to work together to make this ecosystem better."
v0 Ships New Sidebar with Project Grouping
Chats grouped by project with real-time status indicators and hover cards showing website previews and git branch info.
Replit Users Run 100K Daily AI Security Scans
Over 100,000 daily scans with Semgrep make Replit the all-in-one platform for building secure apps with AI.
Replit MCP Beta Gets Full App Management
Create, load, search, and publish apps directly via MCP, enabling software creation on Replit from anywhere.
Google Cloud and Replit Host Agentic Cinema Hackathon
Build multi-step agents using Gemini Enterprise Agent Platform to automate media and entertainment enterprise workflows.
SWE-Bench ProMax: Best Model Scores 41.2%
170 instances, 7 languages, 11.4 files per task. AI coding agents still face significant challenges.
Two Diffusion LM Workshops at NeurIPS 2026
Both in Sydney. Topics: unified formulations, parallel decoding, test-time scaling, hardware-aware serving.
vLLM x NVIDIA Dynamo Meetup at vLLM Conference
Aug 24, San Francisco. Tech talks on serving LLMs at scale plus the inference community in one room.
Jeff Dean Shares Final Google Week at KDD 2026
"It has been quite a week." Jeff Dean presented his Google reflections at KDD before closing his laptop.
MiniMax H3 Speed Boost via Sol Engine
Community-driven Sol Engine significantly speeds up local MiniMax H3 inference, making users feel GPU rich.