Perplexity Brings Its Portable Computer to Windows RTX PCs
Run harnesses, agents, and models entirely on your own machine — no cloud upload required — while keeping frontier cloud models one tap away when the task demands them.
Perplexity's Portable Computer is no longer a macOS-only curiosity. The company is expanding its work with NVIDIA to ship the harness on Windows PCs carrying RTX GPUs, letting developers run agents and models locally against their own files and connected apps without sending tasks to the cloud. When a job outgrows local silicon, the same interface still reaches for frontier cloud models.
The pitch is straightforward: unmetered, private intelligence on hardware people already own. CEO Aravind Srinivas framed the computer as "a workspace of humans and AIs," and NVIDIA echoed that local AI is becoming "more private, accessible, and useful" on Windows. For a category long defined by hosted endpoints, the day's signal is that the agent runtime is moving down the stack and onto the device.
Google Cloud and Inferact Make TPUs a First-Class Citizen in vLLM
The two are partnering to bring native TPU support into vLLM, moving Google's accelerators from an afterthought to a first-class citizen of the serving stack. The goal is faster, better-tuned inference on TPUs for the open-source community.
Perplexity and NVIDIA Expand Local AI to Windows PCs
Perplexity is widening its NVIDIA partnership to deliver fully local, unmetered AI on Windows machines running RTX GPUs and the Perplexity harness — part of a broader push to keep intelligence on-device and private.
vLLM Adds Day-0 Support for Intern-S2-397B
A scientific multimodal mixture-of-experts with 397B total and 17B active parameters arrives with day-zero support in vLLM. Built for long-horizon research, it carries 262K context, MTP-accelerated inference, and strong multimodal, reasoning, coding, and scientific-agent abilities.
SGLang Ships Day-0 Support for Intern-S2-397B
SGLang also lands day-zero support for the 397B multimodal foundation model, spotlighting a new pre-training paradigm that learns directly from raw scientific literature pages — no parsing required — to serve scientific intelligence and long-horizon agents.
"There are two ways AI progress could go very badly. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity."
Sam Altman, OpenAI
GPT-6 Astra Drives End-to-End Testing Inside Codex
A Perplexity engineer is using GPT-6 Astra within Codex to build test harnesses and mock third-party API responses, so teams can verify how the pieces of a system actually work together — end to end — before anything ships.
MiniMax H3 Opens Its Weights as Video Generation Keeps Accelerating
Native stereo audio, multimodal reference control, and an open ecosystem that keeps compounding.
MiniMax is pushing its H3 video model forward on two fronts at once. The model ships with native stereo audio and multimodal reference control, while the open-source community around it — including SGLang and the VDN-H3 pipeline — keeps making the capability faster, cheaper, and easier to build on. On eight B200 GPUs, H3 now generates 14.4 seconds of 768p video in just 9.0 seconds end-to-end, above 2× real-time with no measured quality regression.
Cohere CEO: AI Doomsday Talk Is Too Sci-Fi
Aidan Gomez told Bloomberg the existential-risk debate "veers too far into science fiction" and should stay out of public conversation.
Jensen Huang on Regulation, Open Models, and Infrastructure
At the All-In Summit, Huang argued for sensible regulation, open models, and continued investment in AI infrastructure.
Runway's Drawing App Animates Sketches in Real Time
A new Runway Labs experiment animates your sketch as you draw it, powered by real-time generative world models.
Replit Routines Monitor Production Data on a Schedule
Routines watch production data and hand off anomalies to an Agent that investigates when reasoning is required.
Adobe Firefly Ships a Free AI Image Upscaler
A free online tool raises resolution, clarity, and sharpness while preserving detail across photos and designs.
Vercel Hires the Creator of Google Cloud Run
Steren joins to lead Fluid compute products, with the message that serverless was the last chapter — agents are next.
An AI Camera Helper for Low-Vision Users
Inspired by Be My Eyes, an app lets blind and low-vision users ask AI what is in front of their camera.
FlyOCR Reads PDFs with a Fruit-Fly Connectome
A playful experiment trains a model on the male fruit-fly CNS connectome to parse documents.
Recraft Puts GPT Image 2.5 Flare in the Ring
A side-by-side image-generation comparison across creative scenarios asks readers to pick the winner.
Kling AI Talks Filmmaking at TIFF
At Toronto's film market, Kling AI argued AI makes production more affordable and accessible.
Computers Are Becoming Human-AI Workspaces
Aravind Srinivas sees the computer evolving into a shared workspace for people and AI together.
Mistral's Mensch: Build Your Own Models
Arthur Mensch urges enterprises not to slow down building and owning their own AI models and systems.
Musk: AI Data Centers Lower Electricity Prices
A counterintuitive claim: AI data centers end up reducing consumer electricity prices.
SAS Makes Sparse Attention End-to-End Trainable
A new method injects selector scores into the softmax gate so ranking learns to match predictive impact.
Sakana Namazu Powers a Medical Evidence Finder
The Japanese-specialized model joins a doctor-facing tool that cites and verifies literature.