UK AISI Evaluates Claude Mythos 5 and GPT-5.6 Sol in Landmark Cybersecurity Red-Team Drill
Models attempted tasks with safeguards removed in controlled evaluation; findings set new benchmarks for frontier AI risk assessment.
The UK AI Security Institute has published its latest cybersecurity evaluation of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. In a deliberately stripped-down testing environment where normal safety guardrails were removed, both models were tasked with completing cybersecurity assignments under controlled conditions. The report examines how frontier models behave when external constraints are lifted, providing critical data for regulators and developers navigating the next generation of AI safety frameworks. The findings arrive at a pivotal moment as governments worldwide grapple with how to assess and mitigate risks posed by increasingly capable general-purpose models.
Qwen3.8-Max Ships: Stronger Performance at Lower Cost, Available Now
Alibaba's Qwen team released Qwen3.8-Max, claiming improved benchmark performance alongside reduced usage costs. The model is available immediately for public testing, continuing the rapid cadence of releases from the Qwen family that has reshaped the open-weight landscape this year.
FLUX 3 Lands on Runway with 20-Second Video and Audio Generation
Black Forest Labs' FLUX 3 is now available through Runway, offering text-to-video and image-to-video generation with up to 20 seconds of synchronized audio. The release extends Runway's creative toolchain with one of the most anticipated video models of the season.
I would rather be an optimist and work hard than a pessimist posting about why things won't work. No amount of "it will never work" essays will drive society forward.
@sama on the case for optimism in AI
OpenAI Discloses Two Security Incidents from External Red-Team Evaluations
OpenAI detailed two new incidents that occurred during external cyber evaluations conducted by independent partners. The company outlined how the activity was contained and how it is working with evaluators to strengthen third-party testing protocols. The disclosures reflect a growing emphasis on transparent safety practices across frontier AI labs.
Mistral Unveils Shieldstral: A 3B Open-Weight Safety Classifier for On-Device Deployment
Mistral AI introduced Shieldstral, a 3-billion-parameter multimodal safety classifier released under open weights. The model outperforms competitors up to seven times its size and is designed for on-device content moderation, marking Mistral's entry into the safety infrastructure layer.
MiniMax H3 Tops Video Arena Across Both Text-to-Video and Image-to-Video Tracks
MiniMax's H3 model has claimed the number one spot across both Text-to-Video and Image-to-Video leaderboards on Video Arena, establishing itself as the leading open model for video generation according to community rankings. The achievement underscores the rapid progress in open video generation models this quarter.
NVIDIA Releases Alpamayo 2 Super: Open Reasoning Model for Autonomous Driving
NVIDIA launched Alpamayo 2 Super, now commercially available for robotaxis and autonomous vehicles. The open reasoning model adds 360-degree awareness, high-level driving decisions, and automated reasoning labels, built for complex real-world driving scenarios.
Qwen-Image-3.0-Pro Leaps to Global 5th Place
Alibaba's Qwen-Image-3.0-Pro has surged to fifth in global rankings with a massive leap from the previous generation.
Maestro v1.5.5 Ships with H3; Community Ports It to Untested Hardware
Within 48 hours, the community got H3 running on gaming GPUs and Macs fully offline — hardware MiniMax never tested.
H3 Licensing Clarified: US, EU, UK, and South Korea Approvals Now Available
MiniMax denied claims that H3 cannot be used in certain regions, confirming formal authorization for deployment in the US, EU, UK, and South Korea.
Hy ASR 3.0 Preview: Speech Recognition Powered by Hunyuan's Language Brain
Hy ASR 3.0 leverages the Hunyuan model's language capabilities, delivering improved general recognition, context awareness, multi-scenario robustness, and dialect coverage.
Cursor Open-Sources MoK: MoE Training Megakernel, 2.37x Faster
Mixture-of-Kittens fuses all MoE communication and computation into a single deterministic kernel for NVL72s.
Day-0 Support for Mistral Shieldstral-1.0-3B Lands on vLLM
The 3B safety classifier, based on Ministral-3 with a Pixtral vision encoder, gets immediate inference support for edge deployments.
Seedance 2.5 Previews: AAA Visuals, Realistic Motion, VFX-Grade Control
Higgsfield is teasing Seedance 2.5 with demos spanning cinematic atmosphere, realistic human motion, and video-to-video editing — all arriving soon.
Pixar Co-Founder Edwin Catmull Joins Higgsfield Film Festival Jury
The Turing Award laureate and 5-time Oscar winner will judge the first generation of AI-generated films at the Higgsfield Global Film Festival.
Factory Scales to One Billion Monthly Requests on Vercel
Factory's backend migrated to Vercel using Next.js, Fluid Compute, and Vercel WAF, now handling a billion requests per month without a dedicated security team.
SpecForge v0.3.0: Unified Disaggregated Speculative Decoding Stack
The release adds fully disaggregated online training and support for EAGLE3, DFlash, Domino, and DSpark draft models in a single runtime.
DeepSeek-V4-Flash Becomes Ollama's Fastest-Growing Model Ever
Running at 100+ tokens per second with zero data retention, capacity is scaling in the US and Europe.
Ambient Intelligence Surfaces Design Suggestions Beside Every Frame
Suggestion cards offer different design directions. Selecting one auto-generates a new frame.
Pika Launches API Club: $10/Month, Up to 88% Lower Model Pricing
Pika API Club offers a monthly membership with usage-based pricing that the company claims undercuts aggregators by up to 88%.
LlamaParse Adds Forms Enrichment: W-2 Parsing Returns Structured JSON
Set processing_options.forms to 'enrich' and LlamaParse returns dedicated JSON with field names, values, and checkbox states.
Day-0 Support for Liquid AI's LFM2.5-2.6B on SGLang
The hybrid architecture model with ~34T pre-training tokens serves many concurrent agents on a single GPU.
Pokee-Isaac 28B Gets Day-0 SGLang Support: 10M-Token Context on One GPU
93.3% RULER at full 10M-token context, up to 137K tokens/s prefill on a single B200.