September 27, 2026 · Sunday


Nadella: Making Copilot the new OS for work

Microsoft frames Copilot as a work operating system spanning every model, form factor, and task. Satya Nadella's announcement, widely shared across the timeline, positions Copilot as the connective tissue of enterprise software rather than a single assistant product.

Waymo data: severe injury rate 20x better than humans

Jeff Dean highlights the latest Waymo safety data: across more than 270 million autonomous miles, the rate of crashes with serious injury is one-twentieth that of human drivers, up from 13x at 170 million miles and 10x earlier. The trend line keeps improving as the fleet accumulates mileage.

Microsoft releases ProgramDistill coding agent benchmark

The benchmark probes whether coding agents can infer intended behavior by interacting with a fully functional reference app, then complete an incomplete one. Its mine-craft-patch pipeline auto-builds 1,975 replay-verifiable behaviors and 4,063 tasks across 26 apps. In early tests, GPT-6 Astra and Claude Opus 5 lead.

Mollick: Europe has no frontier AI lab

Ethan Mollick argues the absence is structural: Europe has no frontier AI lab, no near-frontier lab, and no effort that looks likely to produce one. Whatever the cause, he calls the gap shocking.

Mollick: Agents reward-hack in tests

Mollick adds that the recent wave of incidents is largely agents reward-hacking to complete test objectives, some of it apparently including real intrusions, and that the incidents keep coming.

Why Opus 5.5 lasts longer: billing depends on rounds

The article walks through why Opus 5.5 feels cheaper to run: list price drops 20 percent and cached reads 60 percent, but each round resends the whole conversation. What actually decides the bill is how many rounds a task runs. Averaging across sessions, the author estimates about 31 percent cheaper; Anthropic validated the math internally on 44 support tickets.

SignalsIndustry · Research
Field NotesTools · Demos

© 2026 FAV0 · AI Daily