October 10, 2026 · Saturday

If you want to really understand AI RSI, science should be your reference point.

— François Chollet
● Products & Releases10.10 · Briefs

OpenAI math problems become open RL environments

The problems are repackaged as open-source reinforcement-learning environments for direct training and reproduction.

VERCEL

Agents are starting to buy infrastructure

The Vercel CEO says agents buy cloud products and services via CLI more than expected; the capability now extends to domains.

RESEARCH

LeWAM: JEPA imagines state and action

A JEPA for the real world that predicts both what happens next and what to do, with 32.3x faster training.

FINANCE

Frontier models beat human forecasters

New results claim frontier AI models outperform human experts on financial-prediction tasks.

NVIDIA

GR00T robot model hits Hugging Face

Agile One S SSD Pick is a GR00T-based deployment model for SSD carrying tasks.

SEMIANALYSIS

Rubin inference profit per watt up 3.2x

NVIDIA's vLLM Rubin inference boosts profit per gigawatt 3.2x and per-dollar performance up to 10x.

COMMENT

Neel Nanda questions OpenAI's account

The interpretability researcher says the firing story "doesn't add up," laying out two possibilities.

AI tutoring raises scores; ghostwriting fades

A randomized GPT-4o trial found tutoring gains persist a week later, while having AI write for students does not.

Older Gemini matches doctors on ER advice

Gemini 2.5 Pro and Flash advice was rated comparable to doctors in urgent care, with no safety issues spotted.

MICROSOFT

Decision-1: a model that only chooses

A Qwen3.5-9B-based model that judges given options and returns probability scores, live on Microsoft's platform.

DEEPMIND

AlphaProtein Novo designs enzymes

A pipeline for de novo enzyme design, built with Frances Arnold's lab.

OPENAI

ChatGPT dots now start on mobile

You can create your dot straight from the ChatGPT app on iOS and Android.

ALIBABA

Qwen3.8-Max free for a week

Qwen3.8-Max, Qwen3.8-Flash, and Wan3.0 are free on GMI Cloud for one week.

LUMA

Product Shots reuses one image

Place the same product shot into new settings without rebuilding from scratch.

FILM

Moses shot on stage with AI

Jon Erwin used real actors plus AI to carry them into any world the story needed.

VERCEL

v0 for teams ships

See what teammates are working on, jump into their chats, and build together.

LLAMAINDEX

LlamaParse untangles nested tables

18 values correctly assigned under the right headers in Micron's earnings deck.

TOOLING

Animation from code, not video gen

Claude plus Higgsfield Katana replaces After Effects and Blender — pure code.

IDEOGRAM

Ideogram 4.5 claims precision

The team calls 4.5 the most precise image-edit model available.

IDEOGRAM

40 edits, head to head

Nano Banana 2.1 and Ideogram 4.5 each iterated on their own last image, same prompt.

DEEPMIND

Bending the Curve of Discovery

Hassabis and Manyika publish a new essay on AI and science.

PERPLEXITY

pplx-decider tops Decision Bench

94.5% accuracy at the lowest cost on the leaderboard.

PERPLEXITY

93 hours of explainers

Computer made explainer videos for all 722 OpenAI math preprints.

MODELS

Whistle: 16.9MB speech-to-text

A tiny on-device model that rivals Whisper base.

HEALTH

MedDecider goes open-weight

Decision models for medicine reach leading benchmark performance.

INFRA

Open weights net-boost AI demand

Compressed model-layer margins, but expanded usage means more infrastructure.

OPEN SOURCE

Busabase: memory for agents

An MIT-licensed database so agent output stops getting trapped in chat history.

ARC

ARC-AGI-3 hits 59.17%

Yi-Chia Chen leads the ARC Prize 2026 board, ahead of Tufa Labs.

OPENROUTER

Provider token ranking sparks debate

A top-five list by daily tokens draws scrutiny over methodology.

COMMENT

Efficiency beats breakthroughs

Natolambert bets on efficiency gains as the bigger lever for diffusion.

DATA

Datology's curation playbooks

Practical reports on filtering data across pre, post-training, and evals.

RESEARCH

PoolDINO drops 4-16x tokens

RAE reaches comparable quality with far fewer tokens.

COMMENT

A tenth of Opus's cost?

Stas Bekman weighs openai-gpt-6.1-sol for cheaper code quality.

RL

TermGrade for terminal agents

Open-source reinforcement-learning environments, fully open.

DEVELOPER

A blog feature, built by voice

Simon Willison used Codex Desktop voice mode while cooking dinner.

COMMENT

Google's test after Gemini 4

Mollick: a single interface plus orchestrated agents, not fragmented products.

COMMENT

More proofs aren't progress

New goals are needed for what math is trying to do.

ANTHROPIC

Startup perks paused after 3 days

Free Claude Team plans and $1,000 API credits paused as demand exploded.

● Short Takes10.10 · Closing Wire

© 2026 FAV0 · AI Daily