Sakana Fugu Conductor Model Switches to Gemma 4
Sakana AI announced training the Fugu conductor model on Gemma 4, with performance and cost on par with the Qwen version, and modular replacement to suit sovereignty needs. The multi-agent orchestration product treats Fugu as a single foundation model, enabling plug-and-play substitution of both the conductor and the model pool.
Sakana AI has trained the Fugu conductor model on Gemma 4, achieving performance parity with its prior Qwen-based version at equivalent cost. The design makes both the conductor model and the model pool modularly replaceable, allowing organizations to swap components to meet different sovereignty requirements. Future work includes developing a conductor model based on a domestic Japanese model, balancing overseas AI capabilities with national autonomy requirements.
MiniMax H3 Community AMA: Will Remain Open Source
In a Reddit AMA, MiniMax reviewed the H3 architecture and workflow, previewed next release plans, and confirmed it will continue open sourcing.
DeepMind AI Improves Cyclone Prediction Accuracy
Google DeepMind published research in Nature using AI to predict cyclones, buying more time for disaster response. Every hour of lead time counts.
The paper addresses a critical gap in meteorological forecasting. Accurate cyclone prediction directly translates to more hours for evacuations and preparation, potentially saving thousands of lives in vulnerable coastal regions.
Grok Imagine 2.0 Joins Vercel AI Gateway
Grok Imagine Image 2.0 ranks second on Arena AI and is now callable via Vercel AI Gateway, expanding its reach to developers building AI-powered applications.
Musk: AI Agent Traffic Will Vastly Exceed Humans
Musk cited Cloudflare's forecast that internet traffic from AI agents will far exceed human usage. He called it not even a close call.
Denny Zhou: AGI = Transformer + Reasoning
He said all that is left is data and scale; everything else is marginal improvement, sparking debate on the AGI path.
While the AI problems we face seem technically tractable, our incentive structures create an environment where I expect most solutions come after more serious harms.
Nathan Lambert
Open-Weight Model Distillation Rumors: Reasoning Traces Are Key
Rumors suggest models like Kimi, Qwen, and MiniMax are being distilled from larger systems. Good distillation relies on reasoning traces that are normally hidden from users. The speculative narrative has drawn attention to the opaque pipeline of open-weight model development, where model capabilities may reflect hidden training data paths rather than independent architectural innovation.
SmolForge Removes Two Dangerous AI Skills
SmolForge shared lessons from AI coding skills that led to agent overreach. One skill turned a temporary "ship fast" instruction into permanent production authorization; another auto-upgraded simple changes into major overhauls. The team chose to delete both.
Using Codex to Give a Classic Text Adventure a GUI
After the open-sourcing of "A Mind Forever Voyaging" by Steve Meretzky, Ethan Mollick had Codex generate a browser UI, making a modern playable version of this beloved interactive story.
Free Mac Subtitle Tool BaoCut: Local Transcription and Editing in One
BaoCut uses Apple MLX for on-device transcription and speaker diarization, with cloud models off by default. All AI actions generate suggestions and only take effect after confirmation, supporting translation, soft editing, and export. Designed as an agent-first app, it offers a web interface for use within coding agent browsers.
Pika Demonstrates Seedance 2.5 Audio-to-Music-Video
Pika used Seedance 2.5 to generate a music video from audio alone, with standout character and choreography performance that impressed creators.
Is the Codex Benchmark Biased Toward Large Models?
Evaluations show Codex ranks 2nd on GLM 5.2 but falls to 9th on Gemma-4, suggesting the benchmark is overfit to large models.
Hugging Face Gets a New Repo Every 7 Seconds
Hugging Face receives a new repository roughly every 7 seconds, translating to about 12,000 new models and datasets per day.
Grok Build Upgrades to All-in-One Creation Environment
Grok Build now supports generating images and videos with Grok Imagine, becoming an integrated creation platform.
Swyx: Ultracode Is a Major Innovation in Coding Patterns
Anthropic's Ultracode has huge potential in dynamic workflows. One contestant produced solid work with just three prompts.
In US 75% Fear AI; in China 80% Are Excited
A survey shows a high share of Americans fear AI, while China is overwhelmingly optimistic, highlighting a widening narrative gap.
Weekly Top HF Papers: Long-Horizon Agents and Multimodal Generation
Hugging Face's weekly roundup covers long-horizon agents, self-improving reinforcement learning, and multimodal generation.
AI in Academic Journals: Debate Should Focus on Future Capabilities
Ethan Mollick argues the debate focuses too much on current AI capabilities; given long review cycles, it should consider the coming years.
ChatGPT Work and Claude Cowork Should Explain Technical Thinking
Mollick criticizes both products for hiding coding reasoning from non-programmers, calling for PM-like delegation explanations.
LiteParse Extracts Structured PDF Data in Milliseconds
LlamaIndex announced LiteParse can quickly extract structured information including checkboxes, annotations, and vector graphics from PDFs.
Ante Harness Puts V4-Flash Above Grok-4.5
0731 reaches 82.7 on TerminalBench 2.1, matching DeepSeek Harness at lower cost.
Tracking China's AI Research Lineages
Moonshot founded by ex-Google; Yang Zhilin studied under Jie Tang at Zhipu. DeepSeek stands out as unusual in this landscape.
Vercel CEO: Even with AI, You Still Need to Read Code
Rauch argues that skipping code leads to novice, prototype, or debt states. The agent era still requires understanding.
MFU Was a Marketing Metric; Why HFU Never Caught On
Discussion of MFU's origins and HFU's disappearance, reflecting on the motives behind compute utilization metrics.
GPT-4 Completed Training Four Years Ago
Greg Brockman shared a reminder of how far the field has come since GPT-4 finished training.
Codex Helps Read Contract Fine Print to Save Money
Greg Brockman shared a Codex use case: carefully reading contract terms to help users save money.
Halting Endless Agent Delegation in Codex
Mollick demonstrated asking Codex's main model to stop delegating to subagents, avoiding missed critical issues.
Is Concentrating AI Compute Too Risky?
Fleuret questions concentrating enormous value in a single compute point, where one disaster could cause huge losses.