Hunyuan Hy4 Squeezed to 200GB — and It Still Works
The trick isn't just going low; it's deciding where, as calibration data picks each layer's bit width.
Tencent compressed its Hy4-preview model from 1.5 terabytes down to about 200GiB in GGUF format, and reports it still performs well. The new MIX-STQ1_0 scheme uses calibration data to select a per-layer bit width: some layers drop as low as 1.31-bit STQ1_0, while others rise to 2.06-bit IQ2_XXS — all within the same memory budget.
Kimi K3 Arrives in Cursor, Near the Frontier
Moonshot's Kimi K3 is now available in Cursor, running on US-based inference nodes, with CursorBench scores that put it close to frontier models.
MiniMax H3 Goes Open, H3 Max Raises the Bar
MiniMax released H3 with open weights, and the community's fal-led post-train — H3 Max — now resets expectations on both benchmarks and speed.
Watching a popular coding tool lose frontier-model access shows why model resiliency matters: in the future, products will keep working even when several underlying models go offline — they'll just route around them.
— @hardmaru
The Stated Reason Is ToS — the Real Story Is Commoditization
OpenAI's stated reason is a terms-of-service violation. Perhaps. But models are commoditizing, value capture is moving to the application layer, and no one wants to be reduced to a commodity. Competition regulators should be paying close attention.
A Small Window to Build Your Own Intelligence
Most companies now realize they have a narrow window to build their own intelligence — or accept the arbitrary terms set by three private providers.
Claude Code Weekly Limits Rise 25% for Good
From September 14, standard weekly limits climb permanently by 25% on Pro, Max, Team, and seat-based Enterprise plans; the temporary 50% bump stays until then.
ChatGPT Workspaces Sync the Codex GitHub Plugin
Business and Enterprise workspaces can now import the Codex GitHub plugin with admin-managed syncing.
Hy4 Preview Runs Natively in vLLM
Tencent's Hy4 preview can now run directly in vLLM, easing deployment and inference for developers.
GLM 5.3 Rolls Out on Ollama Cloud
GLM 5.3 and GLM 5.3 Flash are fully live on Ollama's cloud — private, fast, US- and Europe-hosted, with zero data retention.
Perplexity Search API Sweeps the Index
The API takes the top three spots on the Artificial Analysis Search Index, extending the quality-cost frontier at about $0.091 per task.
V8.2 Edit Model Gets a Quality Bump
Midjourney updated its V8.2 edit model for better image quality and asks users who hit issues in the last 24 hours to give it another shot.
Vera Rubin Posts 30x Throughput per Megawatt
Are you measuring AI infrastructure on how agents actually run? Measured by NVIDIA on SemiAnalysis's AgentX workload, Vera Rubin NVL72 shows up to 30x better throughput per megawatt than GB300 NVL72.
About 15GW of 2027 AI Compute Can't Switch On
It's harder than finding power: transformers, wiring, liquid cooling, massive chillers, and complex networking all have to be built out too.
LeVJEPA Hits a New Pareto Frontier for Video Pretraining
A stable, efficient end-to-end method that authors say opens many doors for video foundation models.
Gemini Accelerates Real-World Scientific Discovery
Early progress from the Schmidgall lab uses Gemini to speed up real-world scientific discovery.
Autonomous AI Scientists: Promise and Gaps
A new paper maps the potential and remaining gaps of autonomous AI researchers. The big question is how much more advanced models close them.
AI Overviews May Be Doing to Wikipedia What Agents Did to StackExchange
Early evidence suggests Google's AI Overviews are quietly siphoning traffic from Wikipedia.
"Build a Reasoning Model (From Scratch)" Draws Book Clubs
Two book clubs are now discussing it, with live Q&As set for September 3.
Terminal-Bench 4 Ships a Month After v3
The fast follow marks a shift to a continuous-style benchmark, worth studying for high-level improvement patterns.
SGLang Visualizes Kimi K3's Architecture
New tutorials cover Kimi Delta Attention, hybrid attention radix trees, and TP/DP strategies.
fx 0.0.7 Improves MCP Support
The tiny Zig coding agent now handles top MCPs — Context7, Datadog, MongoDB, Linear, Notion, Supabase — with a leaner toolset and improved shell execution.
Infinite Slop: An Endless AI Livestream
levelsio and fal.ai ship a chat-driven AI channel that generates forever — whatever you type becomes the next video.
Gemini Omni 1.1 Flash Renders a Macro Ad in One Shot
Refraction that is notoriously hard to keep stable, made to look effortless.
The Trigger: An Open-Sourced AI Short Film
Eight minutes of action and one impossible hitman mission, made on Cinema Studio 4 for a $1M film festival.