Qwen3.8-27B puts a frontier model on your laptop
Alibaba's compact open model lands on LM Studio, billed as a "laptop-sized, frontier-level jump" that runs entirely on-device.
Alibaba's open-source Qwen3.8-27B is now live on LM Studio, letting developers download, deploy, and run a 27-billion-parameter model on an ordinary laptop. The Qwen team describes the release as a "laptop-sized, frontier-level jump" — a deliberately compact model that keeps the reasoning and coding muscle of far larger systems while fitting into a single workstation.
The launch caps a breakneck week for the model: it already runs on NVIDIA's RTX Spark, has been paired with MediaTek for phones and vehicles, and within hours climbed to fourth place on Hugging Face's all-time most-liked list. On-device AI is no longer a promise — it is shipping to developers today.
Runway ships Seedance 2.5 in 1080p
Sharper detail at higher resolution — early access starts today.
Runway has switched Seedance 2.5 over to 1080p, bringing visibly sharper detail at higher resolution as early access opens. The upgrade is already rippling beyond Runway's own platform, with the model surfacing on Hedra and sparking a fresh wave of community experiments.
DeepSeek-V4-Pro cloud lands on Ollama
DeepSeek-V4-Pro-0813 is now fully rolled out on Ollama Cloud and bundled into Pro and Max subscriptions. It is hosted in the United States with zero data retention, and can be pulled with a single command.
vLLM adds dynamic draft length for DSpark
Six weeks ago, serving DSpark in vLLM meant fixing a draft length for your traffic and living with it. Now vLLM decides how much of each draft to verify at every step — and the first token of a 7-token draft survives on DeepSeek-V4-Pro-0813.
Grok 4.6 tops news reliability benchmark
Grok 4.6 ranks first on RuntimeWire's Newsroom Reliability v0.2 with a score of 0.79, beating GPT-5.6 Sol and Claude Opus.
Grok 4.6 joins Microsoft Copilot
Elon Musk confirmed that Grok 4.6 is now available inside Microsoft Copilot, expanding the model's reach into a major productivity surface.
Grok 4.6 clears the Gauntlet
Musk announced that Grok 4.6 has run "The Gauntlet," a demanding agentic evaluation, signaling the model's push into long-horizon agent work.
On regulation: either concentrate AI in the hands of a few — or build far more careful institutional arrangements.
— Dario Amodei, CEO of Anthropic
Pika launches four audio models
Video dubbing, TTS, music, and SFX models arrive at half to one-twentieth the price of rivals, targeting cost-effective audio generation.
Qwen3.8-27B teams with MediaTek
Alibaba is putting Qwen3.8-27B into smartphones and vehicles through a MediaTek partnership for on-device AI.
Qwen3.8-27B flies on RTX Spark
The model can be downloaded and deployed on NVIDIA's RTX Spark device, with official builds now available.
DeepSeek-V4-Pro-0813 arrives on Novita
Novita now supports the model on Hugging Face with a 1M-token context window built for reasoning and coding.
Ideogram updates its AI image editor
A new editor adds object removal, local editing, and painting features.
Qwen3.8 open weights released
Qwen3.8-27B's weights are now open, positioned as a native multimodal dense model with multiple local deployment options.
Arcee open-sources agent harness Nac
Nac is a harness for long-running tasks, launched alongside the Arcee open models API beta.
Gemini 3.7 Flash hits the Gemini app
DeepMind's CEO confirmed Gemini 3.7 Flash is now available in the Gemini app.
NVIDIA releases MOPD expert models
NVIDIA's expert models for MOPD make related research far more accessible.
llama.cpp runs Qwen3.8-27B in one line
The GGUF build runs via llama serve with draft-mtp speculative decoding.
Zhipu invites teams to evaluate GLM-5.3
Zhipu is inviting cybersecurity organizations and researchers to jointly assess GLM-5.3's safety.
"Prompting is going away. Delete everything, keep Graph."
— Andrej Karpathy
Codex helps a kernel go 232x faster
A Hacker News post shows Codex and other AI tools accelerating kernel code by 232x. Commenters used DeepSeek v4 and Opus 5 to optimize video codecs and write NEON kernels.
GLM-5.3: same base, big post-training gains
GLM-5.3 and GLM-5.2 share a 756B base; post-training delivers the gains. DeepSeek V4 Flash does it with 304B parameters, and Luna is estimated around 500B.
Will AI erase adoption friction?
Ethan Mollick flags the big question: if systems keep improving, the usual frictions of technology adoption may simply disappear — though the direction is still unclear.
Twitter's algorithm gains a 14-day video recall
A new open algorithm adds a 14-day SID semantic recall window for videos, and folds follows and mutual relationships into recommendations, giving quality videos a second distribution.
DSH vs Pi: almost everything is a plugin
Pi aims for a minimal core plus user assembly; DSH makes nearly all capabilities pluggable with official combos, built on the Cordis skeleton.
Mercor's secret project hires 26,000+ workers
Reddit users found AI data firm Mercor's large project has hired more than 26,000 contractors under confidentiality, sparking discussion.
MiniMax H3 shows off creators in SF
At a San Francisco event, creators demoed H3 for commercials and music videos, with a technical walkthrough and fireside chat.
The Gauntlet Loop prompting method
A blog details the Gauntlet Loop: set clear standards, split the task, don't let the builder self-grade, and keep iterating — for code, design, writing, and research.
Gemini 3.1 Pro is good enough
A Google researcher argues mainstream models like Gemini 3.1 Pro cover 99% of everyday work — not everyone needs the strongest coding model.
Compression is RNN, recursion is attention
An analogy: context compression resembles an RNN's fixed state, while recursive reasoning models resemble attention's full-context reprocessing.
Compiler-agent co-design is next
ML compilation experts say co-designing compilers and agents is the next frontier for kernel optimization.
Codex revives a 1987 Infocom game
A blogger used Codex to turn 1987's "Nord and Bert" into a graphical version, keeping the puzzles while improving UX and AI-translating some.
Indie devs face harsher AI scrutiny
Ethan Mollick observes that resource-strapped indie game developers face harsher public backlash for using AI than big studios.
Training keyframe LoRAs with two-frame animation
A blog shares a method to train keyframe animation adapters on pretrained video models, controlling pacing with a few keyframes.
Inside DSH's mental model
An updated article argues DSH's internal mental model is more complex than expected, yet simple from a plugin developer's view.
What DSH's "everything is a plugin" means
A discussion contrasts DSH's plugin architecture with Webpack and hook patterns, arguing it pluginizes nearly all capabilities.
Frontier labs' ARR tops Windows and Office
ARK analysts estimate frontier AI labs' annualized revenue surpassed Microsoft's Windows and Office revenue combined by late June.
Self-driving spreads across the Bay Area
The San Francisco Chronicle describes a "quiet revolution" of privately owned cars now driving themselves on Bay Area roads.
GPT-5.6 builds games in Unity
Unity CLI plus GPT-5.6 SOL Ultra and Luna Max drive game development.
Codex becomes a video editor
OpenAI staff use Codex to produce explainer videos.
Qwen3.8 is HF's 4th most-loved
Hours after release, it hit fourth all-time on Hugging Face.
Seedance 2.5 lands on Hedra
The model is also available on Hedra.
SGLang runs Qwen3.8
The team is collecting feedback and shipping improvements.
ChatGPT adds booking search
Reservation queries can now be completed directly.
Perplexity pushes Agent API
For frontier web search, it's hard to beat, says the CEO.
Seeking the next Disney
Rauch: AI lets everyone create — now fund it.
Grok Bot called best agent
A week of use earns it a top pick from one user.
Grok Build makes mini-games
One-line prompts build games in minutes.
Qwen3.8 builds a playable FPS
A user generated a first-person shooter with mouse-look.
Qwen3.8 over-thinks by default
At "extra high" reasoning it's a chronic over-thinker.
A Vinge-style singularity
We're in one by Vinge's definition, not von Neumann's.
Experts should use AI video
To show dinosaurs, Io, and Ancient Rome directly.
AI as media's computing model
Valenzuela: AI is media's new computing paradigm.