August 27, 2026 · Thursday

Alibaba ships Qwen3.8-Flash, a 125B open-weight preview of Qwen4

A multimodal Mixture-of-Experts with a 51B N-gram table previews the next Qwen architecture — and it is already priced for production at $0.16 per million input tokens.

Qwen3.8-Flash launch graphic
Qwen3.8-Flash is an early preview of the Qwen4 architecture, released with open weights.

Qwen3.8-Flash is Alibaba's first look at the Qwen4 architecture, and it ships open-weight. The model combines 125B total parameters with a separate 51B N-gram table, producing a multimodal MoE that handles text and vision in a single system. The production build arrives on QwenCloud at $0.16 per million input tokens and $0.47 per million output tokens — a price aimed squarely at high-volume agent workloads. Inference partners are already lining up: both vLLM and SGLang announced day-0 support, with the architecture verified across NVIDIA and AMD GPUs.

America is LLM-pilled. China is world-model-pilled.

— @c_valenzuelab
Gemini 3.5 Transcribe model visual
Google's latest speech-to-text model is built for smart, precise transcription.

Gemini 3.5 Transcribe chases precision, not just words

Google's new speech-to-text model is engineered for understanding, not merely word-matching. It auto-detects more than 85 languages out of the box, adapts to custom vocabulary for specialized jargon, and can follow multiple speakers within a single call. The API is available now in Google AI Studio and the Gemini API, with function calling built in for agent-style workflows.

Claude Code feedback report

Claude Code drafts its own feedback reports

When something fails, or Claude notices it made a mistake, it writes up the report itself. Developers review, edit, and approve the feedback before it is sent — a small feature that shifts error reporting from manual drudgery to a supervised agent loop.

AGENTS & TOOLINGAUG 27
BRIEFSAUG 27

© 2026 FAV0 · AI Daily