Anthropic and Accenture build independent frontier-AI evaluation
Both sides expect to invest at least $1 billion each over the next five years.
The partnership extends Anthropic's recent commitment to embed evaluators inside the company. Independent, third-party assessment of frontier models has become a flashpoint as labs race to ship more capable systems, and Anthropic now says it will jointly build capacity to test those systems outside its own benchmarks. The move acknowledges a structural problem: the people who build frontier AI are rarely the right ones to judge its risks.
Concentration of power in a few labs is the biggest risk in AI.
Clément Delangue · Hugging Face CEO
MiniMax open-sources its terminal coding agent
MiniMax Code, or mcode, understands a project, edits code, and runs tests from the terminal. It supports a MiniMax account or bring-your-own-key mode, plus search, plugins, and multimodal tools. Install scripts cover macOS, Linux, WSL, and Windows.
JEPA-Anything learns predictive models across worlds
The paper proposes orthogonal predictive decomposition to extend JEPA, decomposing latent targets into complementary factors. Evaluated across vision, biology, clinical data, control, molecular dynamics, physical fields, and weather, it improves most dynamics tasks and cuts single-intervention prediction error in one control benchmark.
Jev sets the fastest AI Gateway adoption record
Vercel reports that Jev reached roughly 13% of teams on its first day — double the GPT-5.6 family and six times Fable 5.1 at launch. The speed of uptake makes it the fastest-adopted model in AI Gateway history.
What actually makes a coding-agent harness work
An empirical study fixes the execution loop and varies only planning, action space, and context management, testing 176 configurations across four models on SWE-Bench Verified and Terminal-Bench 2.1. Tighter context budgets reward better context management, and rule-based pruning followed by LLM summarization proves most efficient.
Meta SAM 3.1 lands on Model API
Meta's Segment Anything Model 3.1 is now available on the Meta Model API, giving developers a fast and lightweight model for segmentation.
Gemini breached three firms in a security test
A report says Google's Gemini model hacked three companies as part of a May cybersecurity evaluation run by a testing company.
Claude Code now supports AGENTS.md
Starting with Claude Code 2.1.277, the agent reads AGENTS.md when no CLAUDE.md is present, unifying agent-rule configuration.
ChatGPT multi-account support reaches plugins
OpenAI adds multi-account support across most plugins, letting users switch between personal, work, and side-project accounts.
ChatGPT desktop supports Chrome extensions
OpenAI brings everyday browser extensions into the ChatGPT desktop app, launching support for Chrome extensions.
Mistral denies its systems were breached
After a thorough investigation into an unauthorized-access claim, Mistral says it found no evidence to support it and confirmed its systems were not compromised.
Three paths to faster video attention
MiniMax visualizes three acceleration paths for video attention and highlights VC-Attention's training-free low-bit acceleration.
Runway Ruby supports alpha channels
Runway adds alpha-channel support to Ruby, converting footage to HDR while keeping transparency fully intact in a single step.
Runway ships nine updates in 18 days
Runway's CEO lists recent launches: Ruby, Dev MCP, Solaris, GWM Worlds 2, Team Plan, Adobe plugins, model licensing, frame-rate enhancement, and alpha channels.
Replit races from $2.8M toward $1B ARR
A repost says Replit had $2.8 million in annual revenue two years ago and now targets more than $1 billion ARR by the end of this year.
Jev's rise rides the cost zeitgeist
Vercel's CEO calls Jev a great product, but argues its rapid adoption is also downstream of the "AI is too expensive and slow" mood, as people rush to optimize and place AI in more places.
Musk predicts AI will double US GDP growth
Elon Musk speculates AI could lift US GDP growth from about 2% to around 4% next year, and perhaps even more.
WSJ: Hugging Face hack was exaggerated
A reposted WSJ opinion argues the "rogue hive mind" framing of AI agents in the Hugging Face incident was overstated.
Sakana Fugu Max arrives on Sakana Chat
Sakana AI makes Fugu Max available on Sakana Chat, free for everyone to use.
Closed models remain the risk iceberg
Nat Lambert argues closed models are still the tip of the AI-risk iceberg — easier to start with, more capable, and shipped with leaky safeguards — while fine-tuning open models for attacks is harder.
Researchers back independent AI evaluation
Nat Lambert signs a push for independent evaluation and oversight of frontier AI, naming technical capability and independence as the biggest bottlenecks.
NVIDIA SoL-Pi lands on HF Papers
SoL-Pi proposes recursively scaling automated research with a token-efficient agent harness, now on Hugging Face paper pages.
NetEase Youdao open-sources Confucius4-R2T2
The 1.7B streaming ASR model outputs not just text but executable actions, giving users something they can act on.
Pretraining data, not verifiability, drives math
A reposted blog argues LLMs are unusually strong at math and coding mainly because of pretraining data, not task verifiability.
ChatGPT customization lands next week
An OpenAI developer teases a customization feature arriving next week, possibly tied to custom instructions or model configuration.
OpenArt Arena debuts a creative leaderboard
OpenArt launches a global leaderboard for creative intelligence to compare how AI models generate.
Kernels lets you swap model kernels
Hugging Face Kernels lets users choose which kernel a model runs on, without rewriting the model.
Papers with Code gets an MCP server
Niels Rogge introduces a Papers with Code MCP server and uses Claude Code to research model architectures.
Runway unifies credit pricing
Purchased credits now work across the web app and Runway Dev, alongside new offers.
Custom connectors and audit logs
Replit adds custom connectors to reach external APIs and 65+ audit events for enterprise admins.
LlamaParse opens a parsing series
The first post uses the US EIA Short-Term Energy Outlook to show the stakes of getting parsing wrong.
Jev auto-routes GenAI models
Higgsfield uses Jev to evaluate prompts and pick cost-effective models, balancing speed and quality.
Vidu S2 switches styles in real time
Vidu S2-Editing swaps style, clothing, subject, and background instantly, with invite-code trials.
The cost of a frame has collapsed
Runway's CEO marvels that one person can now make a personal-story video at low cost, unimaginable a few years ago.
Runway is big in Japan without a local entity
A report says Runway has no local subsidiary or country manager, yet Japan is its third-largest market.
CEOs can't delegate AI transformation
Kai-Fu Lee and Semafor discuss why companies need clear ownership of AI transformation.
Open-source AI as the risk antidote
Hugging Face's CEO says most risk is concentrated in a few strong labs, and open-source AI is the mitigation.
Jev achieves generative UI
Vercel's CEO says Jev's generative UI has been implemented externally.
muse tops the App Store
A week after launch, the AI app muse is No. 1 on the App Store amid strong user enthusiasm.
lm15 is a lightweight litellm alternative
The DSPy team launches lm15, a lighter alternative to litellm supporting Foundry-hosted models.
XGEN unveils generative world simulation
XGEN announces a technical prototype for a generative world simulation system.
Jason Wei speaks at Stanford AI Club
A 30-minute talk on three mental frameworks for understanding the AI landscape, from the researcher behind o1 and chain-of-thought prompting.
Fed up with Photoshop, he built his own editor
Compositor, a Mac Photoshop alternative, is open-sourced with Claude in the commit history — likely AI-assisted.
ChatGPT Pro 20x plan reopens
The plan is available again for users who held ChatGPT Pro 20x in the past 30 days before cancellation.