GPT-5.6 Sol Triples ARC-AGI-3 Score After Fixing Harness Bug That Blocked Memory
OpenAI found that two overlooked API settings had prevented the model from remembering what it learned during benchmark evaluation, artificially suppressing results on the prominent reasoning test.
The ARC-AGI-3 benchmark of 2D puzzle games had long seemed like an Achilles' heel for GPT-5.6 Sol, puzzling researchers since the model had proven capable of solving open problems in mathematics. OpenAI's investigation revealed the root cause was not a capability gap but a harness constraint: the evaluation framework was not configured to let the model retain what it had learned across successive problems. Simply enabling two API settings tripled scores. The episode underscores how sensitive frontier-model evaluation remains to test-harness configuration, and raises questions about how many other benchmark results may reflect infrastructure quirks rather than genuine capability limits.
GPT-5.6 Sol Cracks Long-Standing Open Problem in Probability Theory
Greg Brockman confirmed that GPT-5.6 Sol solved another longstanding open problem, this time in probability theory. The result follows earlier successes in pure mathematics and reinforces the model's emerging role as a research tool for formal reasoning. Brockman's brief post, with characteristic understatement, read simply: "5.6 for solving another longstanding open problem, this time in probability."
OpenAI Opens Frontier Models to 10,000 Researchers, Plans Expansion to 100,000 by 2027
The ChatGPT for Academic Researchers program launches with free access for scientists, mathematicians, and engineers.
OpenAI launched an ambitious academic access program designed to put frontier models in the hands of working researchers without cost barriers. The initial cohort of 10,000 scientists, mathematicians, and engineers will receive free access, with a stated plan to scale that number to 100,000 through 2027. The program, branded as ChatGPT for Academic Researchers, is explicitly framed as a tool to accelerate discovery across disciplines rather than as a product play, signaling OpenAI's effort to embed its models deeply into the scientific workflow while the competitive window remains open.
OpenAI Quietly Open-Sources Codex Security CLI, HN Discovers It First
A new open-source command-line tool from OpenAI escaped official announcement only to be discovered by the Hacker News community within hours. Codex Security CLI scans code repositories for vulnerabilities, tracks security findings across runs, verifies fixes, and can be integrated directly into CI/CD pipelines. The early release suggests OpenAI is testing community adoption patterns for developer tools distributed outside its commercial product surface, and the strong organic HN reception validates the approach.
A significant portion of current AI discourse is less about technological capabilities and more about frontier lab employees navigating their own self-esteem and sense of identity.
François Chollet, Google AI Researcher
Nathan Lambert Warns AI's 996 and 002 Work Culture Is Unsustainable
In a candid essay, Interconnects AI author Nathan Lambert described the industry's punishing work schedules, which range from 996 shifts to so-called 002 patterns where engineers work from midnight to midnight with only two hours of rest. Drawing parallels to elite athletic training, Lambert argued that prolonged sleep deprivation erodes creativity and judgment, two faculties AI research depends on most. The essay also noted that new laboratories are structurally unable to catch up with established teams, not for lack of talent but because the accumulated advantages of internal tooling, data resources, and institutional knowledge compound over time.
Inside Olmo 3 Post-Training: Lambert and Geng on the Messy Reality of DPO
In a rare deep-dive lecture and podcast combination, Nathan Lambert and Scott Geng unpacked the full Olmo 3 post-training pipeline, focusing on the gap between research-paper DPO results and what actually works in frontier-model training. The discussion covers how a research idea survives the journey from whiteboard to production model, and why direct preference optimization is far messier in practice than the literature suggests.
MiniMax and Fireworks AI Jointly Open-Source M3 Inference Kernel for NVIDIA Blackwell
MiniMax and Fireworks AI released open-source inference kernels optimized for NVIDIA's Blackwell architecture, including dense FlashAttention implementations and a sparse KV-outer prefill backend. The kernel automatically selects the optimal backend based on the GQA head ratio, falling back from KV-outer only when the ratio drops below 8 to 1.
Tencent Hunyuan Open-Sources AngelSpec, Hitting 2× Speedup on Hy3-A21B
Tencent Hunyuan released AngelSpec, an end-to-end speculative decoding framework that supports both training and deployment in a unified pipeline. On the Hy3-A21B model, AngelSpec achieved between 1.98 and 2.40 times speedup over autoregressive decoding across concurrency levels ranging from 4 to 64, with throughput improvements of over 10 percent.
Dream-Cubed Trains 3D Diffusion Models on Billions of Minecraft Blocks
A research team from Sakana AI published Dream-Cubed, a project that trains both discrete and continuous 3D diffusion models directly on 32-cube resolution Minecraft terrain, using a dataset of billions of blocks. The discrete masked diffusion variant supports conditional generation by biome, inpainting, outpainting, and user-defined block conditioning, without compressing spatial dimensions, preserving fine-grained creative control for players and world builders.
South Korea's SKT Releases 688B Sparse MoE Model A.X K2 on Hugging Face
SKT released A.X K2, a large-scale sparse mixture-of-experts model with 688 billion total parameters and 33 billion active parameters, to Hugging Face for open access. With this release, A.X K2 becomes the largest open-source large language model originating from South Korea, advancing the nation's push for open-science AI development.
vLLM Hits 464 tok/s Single-Batch Decode on Kimi K3 with DSpark and 4 × GB300
The vLLM project set a new peak for single-batch decode throughput under low-entropy reasoning workloads, reaching 464 tokens per second on Kimi K3 using DSpark across four GB300 accelerators. The benchmark is fully reproducible using the public container image vllm/vllm-openai:kimi-k3, and marks a significant milestone for production inference of large reasoning models.
OpenAI Developer Team Uses GPT-5.6 Sol to Optimize Codex Infrastructure
The OpenAIDevs team applied GPT-5.6 Sol within Codex to self-optimize the platform's own infrastructure and performance. The improvements compound across the inference pipeline and agent loop, confirming that the self-optimization capability demonstrated on core serving infrastructure generalizes to developer-facing systems as well.
Grok Now Integrated into Microsoft Copilot
Elon Musk announced that Grok has officially entered the Microsoft Copilot toolset, marking a significant cross-platform deployment for the SpaceXAI assistant.
Cursor Code Editor Launches on iPad
Cursor officially released an iPad version of its AI-powered code editor, offering developers a larger viewport for working with coding agents on the go.
v0 Adds Open-Weight Model Support, Starting with Kimi K3
Vercel's v0 platform now supports open-weight models, with Kimi K3 as the first integration, enabling one-click full-stack web app generation and deployment.
Replit Debuts Model Selector with Open-Weight Support
Replit's new Model Selector lets users choose the best model for each task, spanning intelligence, cost, and speed considerations across GPT, Claude, K3, and DeepSeek families.
AI Community Signs Open Letter for Coordinated Development Pace
Multiple researchers including Neel Nanda signed an open letter arguing that AI progress is too fast and requires coordinated pacing, with a deceleration option kept available.
RARG Introduces Relevance-Aware Grep for Agentic Search
A new paper proposes the Relevance-Aware RipGrep Search Agent, converting relevance into execution priors for coarse-to-fine document traversal in reasoning-intensive retrieval.
HiFi-UMI: 3mm Precision Robot Policies Without Teleoperation Data
HiFi-UMI achieves deployable manipulation policies using only high-fidelity head-mounted demonstration data, matching or exceeding teleoperation baselines on precision insertion tasks.
Elon Musk Claims Grok Voice Ranks First in Agentic Performance
SpaceXAI's Grok Voice has reportedly topped agentic performance benchmarks, surpassing voice models from OpenAI, Google, Alibaba, and others.
Replit CEO Maps Current AI Price-Performance Landscape
Amjad Masad offered a concise comparison: Kimi K3 delivers frontier capability at a quarter of the price, DeepSeek V4 Flash is the low-cost workhorse, and GPT and Anthropic models round out the premium tier.
Production API for Kimi K3 via vLLM Stack
Baseten provides a fast, reliable production API for Kimi K3 with vLLM handling the inference stack. Users can call the 2.8-trillion-parameter model without managing infrastructure.
Serverless Kimi K3 Deployments with Auto-Scaling
Modal announced day-one support for Kimi K3 deployment, exposing vLLM-powered production endpoints that scale automatically with traffic.
Full 2.8T Model with 1M-Token Context
DigitalOcean, in collaboration with vLLM, offers straightforward deployment and scaling of Kimi K3 on familiar infrastructure, with full parameter access and 1M-token context window at launch.
Day-One Support Across All GPU Generations
vLLM serves Kimi K3 across Grace Blackwell, Blackwell, Hopper, and NVL72 systems, integrating with Dynamo to keep the model fast and efficient at production scale.
ROCm Support Available from Release Day
vLLM and AMD partnered to bring Kimi K3 to ROCm at launch, enabling teams to serve the full model on AMD Instinct hardware with broader performance tuning in progress.
Batch Parsing for Up to 10,000 Files
LlamaParse added UI-driven batch parsing that handles up to 10,000 files in a single run, eliminating the need for API scripts and making document extraction accessible without coding.
Transcribe Integrates Superwhisper, Adds Offline Mode
Cohere Transcribe launched in Superwhisper with push-to-talk instant transcription, offline support, cross-app integration, and specialist vocabulary recall for professional workflows.
OAuth Sign-In Simplifies Community Onboarding
Hugging Face introduced a Sign-In OAuth button for third-party websites and apps, enabling community members to share email, create repos, store data in Buckets, and launch GPU-backed Jobs.
Vidu S1 Launches AI Game Companion with Real-Time Voice Interaction
Vidu's new S1 model introduces an AI gaming companion capable of real-time voice conversations and emotional interactions with pets, original characters, and NPCs. The companion uses expressive voices and natural dialogue, with the video demo generated entirely by Vidu S1 and Vidu Q3 using AI.
Pika Officially Confirms Seedance 2.5 Video Model Is Coming Soon
Pika Labs confirmed that Seedance 2.5 will arrive shortly on the platform, addressing recent rumors and signaling another iteration in the rapidly evolving AI video generation space. The announcement follows competitor Hailuo's release of the MiniMax H3 video model and the broader industry acceleration toward production-quality generative video.
Vidu and MachGen Launch Q3 Turbo API for Second-Level Video Generation
Vidu partnered with MachGen to offer the Q3 Turbo model through a commercial API, claiming the fastest AI video inference available with second-level generation at full quality. The service is currently available for free trial, with the partnership signaling growing demand for production-grade, low-latency video generation that matches offline quality.
Z.AI's GLM Model Connects to Pi Coding Agent via API Key
Z.AI published a setup guide for connecting its GLM model to Pi Coding Agent, enabling code generation, debugging, and project analysis directly from the terminal. The integration supports the GLM Coding Plan and works through a standard Z.AI API key, lowering the barrier for developers who prefer terminal-native AI coding workflows.