September 14, 2026 · Monday







Frontier Models & DeepSeek09.14
DeepSeek

V4.1 Flash ships without a base model

Arohan notes that this time DeepSeek did not release the base model alongside V4.1 Flash, breaking with its usual practice.

DeepSeek

V4.1 Flash is a long-horizon agent, not a GLM rival

Dismissing “a bit behind GLM 5.3 Flash” takes, one observer calls V4.1 Flash a true long-horizon agent for open-ended tasks — still half-baked as a product.

Benchmarks

What will V4.1 score on ARC-AGI-2?

A prediction floating around the feed: 75–78 percent “sounds about fair” for DeepSeek’s newest model on ARC-AGI-2.

DeepSeek

DeepSeek’s mission: test-time continual learning

A mission statement from DeepSeek’s Shengding Hu describes a “straight shot to test-time parametric continual learning,” not RSI and not harness-level evolution.

Policy

A practical case against banning open source

Open source should not be banned, one practitioner argues — it sounds dystopian and anti-freedom — even as doubts about its economics persist.

Alignment

Why models scheme: it’s in the pretraining

Scheming behavior on message boards is learned from pretraining itself, the argument goes; better alignment means better training recipes.

Product

Willison puts Astra’s running routes to the test

Simon Willison asked ChatGPT Work and GPT-6 Astra to design 5K and 10K loop routes from his home using OSM data — he got maps and GPX files, but the model’s code stayed invisible.

Product

Talking to a game that builds itself

A developer’s GPT-Live-1 playtest lets you talk to a game to generate its world in real time.


Signals & Short Takes09.14

© 2026 FAV0 · AI Daily